Electronic device and method for providing visual object including translated text

WO2026177333A1PCT designated stage Publication Date: 2026-08-27SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/022358
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-12
Filing Date
2025-12-19
Publication Date
2026-08-27

Smart Images

  • Figure KR2025022358_27082026_PF_FP_ABST
    Figure KR2025022358_27082026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an electronic device and a method for providing a visual object including translated text. A method by which an electronic device provides a translation result for text in a preview image comprises the operations of: acquiring first translated text by translating first text in a preview image; generating a first visual object including the first translated text on the basis of the preview image, the first text, and the first translated text; displaying the preview image and the first visual object, wherein the first visual object is overlaid on the first text in the preview image; and maintaining the first visual object on the preview image when it is not the case that at least a part of the first text has disappeared due to a preset condition.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method for providing a visual object containing translated text

[0001] The present disclosure relates to an electronic device and method for providing a visual object including translated text.

[0002] With the recent advancements in mobile and smart devices, users are increasingly utilizing technologies that leverage cameras to acquire and use various types of information in real time. In particular, there is a growing demand for features that allow for the instant translation and understanding of documents written in foreign languages, leading to the development of real-time camera translation technology. Furthermore, conventional methods typically involve simply obscuring the area containing the original text or displaying the translated text in a fixed format over the original text.

[0003] However, when translating original text within a preview image, if the original text moves or disappears due to the movement of objects within the preview image, there is a need for a technology that can effectively maintain or remove the translation result from the screen while maintaining the visibility of the translation result, taking into account the user's intent and / or the user's attributes.

[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.

[0005] According to one embodiment, an electronic device (1000, 1401) may provide a method for providing a translation result for text in a preview image, comprising: an operation (210) of acquiring a preview image through a camera (1450); an operation (220) of acquiring a first translated text by translating a first text in the preview image; an operation (230) of creating a first visual object including the first translated text based on the preview image, the first text, and the first translated text; an operation (240) of displaying the preview image and the first visual object, wherein the first visual object is overlaid on the first text in the preview image; an operation (250) of identifying whether at least a part of the first text has disappeared in the preview image according to a preset condition; and an operation (270) of maintaining the first visual object on the preview image when at least a part of the first text has not disappeared according to the preset condition.

[0006] According to one embodiment, a camera (1480); a display (1460); a memory (1430) for storing commands; and at least one processor (1420); wherein the instructions, when executed individually or collectively by the at least one processor, cause the electronic device to: acquire a preview image (210) through the camera, acquire a first translated text (220) by translating a first text within the preview image, create a first visual object (230) including the first translated text based on the preview image, the first text, and the first translated text, and display the preview image and the first visual object (240), wherein the first visual object is overlaid on the first text within the preview image, identify (250) whether at least part of the first text has disappeared within the preview image according to a preset condition, and if at least part of the first text has not disappeared according to the preset condition, maintain (270) the first visual object on the preview image, thereby providing a translation result for the text within the preview image. An electronic device (1000, 1401) that provides translation results may be provided.

[0007] According to one embodiment, a computer-readable recording medium may be provided that records a program for executing a method comprising: acquiring a preview image obtained through a camera; acquiring a first translated text by translating a first text within the preview image; generating a first visual object including the first translated text based on the preview image, the first text, and the first translated text; displaying the preview image and the first visual object, wherein the first visual object is overlaid on the first text within the preview image; identifying whether at least a portion of the first text has disappeared within the preview image according to a preset condition; and maintaining the first visual object on the preview image when at least a portion of the first text has not disappeared according to the preset condition.

[0008] The technical problems to be solved in this document are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this invention belongs from the description below.

[0009] FIG. 1 is a diagram showing an overview of an electronic device according to one embodiment providing a translation result for text on an image.

[0010] FIG. 2 is a flowchart of a method in which an electronic device according to one embodiment provides a translation result on an image.

[0011] FIG. 3a is a diagram illustrating an example in which a visual object based on a preview image, a first text, and a first translation text is obtained through an artificial intelligence model according to one embodiment.

[0012] FIG. 3b is a diagram illustrating an example in which a visual object is obtained through an artificial intelligence model by reflecting the user's visual information according to one embodiment.

[0013] FIG. 3c is a diagram illustrating an example in which a visual object including a first translation text and additional information is obtained through an artificial intelligence model according to one embodiment.

[0014] FIG. 4a is a diagram showing an example in which a visual object generated from an artificial intelligence model according to one embodiment is displayed on a preview image.

[0015] FIG. 4b is a diagram showing an example in which a visual object generated from an artificial intelligence model based on user visual characteristic information according to one embodiment is displayed on a preview image.

[0016] FIG. 4c is a drawing showing an example of a visual object including a translation result for text in a preview image and images in a preview image according to one embodiment.

[0017] FIG. 4d is a drawing showing an example of a visual object including translation results and additional information for text within a preview image according to one embodiment.

[0018] FIG. 4e is a drawing showing an example in which a visual object is displayed when the electronic device according to one embodiment is a wearable electronic device.

[0019] FIG. 5 is a flowchart of a method for determining whether an electronic device according to one embodiment maintains the display of a first visual object when the first text, which is the subject of translation, disappears from the preview image.

[0020] FIG. 6a is a drawing showing an example in which a first visual object is displayed when the first text is obscured by another object as the first object containing the first text moves according to one embodiment.

[0021] FIG. 6b is a drawing showing an example of stopping the display of a first visual object when the first text disappears from the preview image as the shooting direction of the electronic device changes.

[0022] FIG. 7 is a flowchart of a method in which an electronic device according to one embodiment predicts whether a first object containing a first text disappears due to another object and performs a display setting for a first visual object.

[0023] FIG. 8 is a drawing showing an example in which an electronic device according to one embodiment changes and displays the attributes of a first visual object.

[0024] FIG. 9 is a flowchart of a method for displaying visual objects when a plurality of objects within a preview image overlap, according to one embodiment of an electronic device.

[0025] FIG. 10 is a drawing showing an example of visual objects being displayed when a plurality of objects within a preview image overlap according to one embodiment.

[0026] FIG. 11 is a drawing showing an example of a visual object fixed to the top layer being moved according to one embodiment.

[0027] FIG. 12 is a drawing showing an example in which the transparency of a visual object is controlled according to one embodiment.

[0028] FIG. 13 is a drawing showing an example of fixing a visual object to an upper layer according to one embodiment.

[0029] FIG. 14 is a block diagram of an electronic device (1401) in a network environment (1400) according to various embodiments.

[0030] FIG. 15 is a drawing showing a system including a generative artificial intelligence model according to one embodiment.

[0031] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.

[0032] Embodiments of the present disclosure are described below in detail with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present disclosure in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.

[0033] The terms used in this disclosure are described in their current, general form considering the functions mentioned herein; however, they may refer to various other terms depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Accordingly, the terms used in this disclosure should not be interpreted solely by their names, but should be interpreted based on the meaning of the terms and the overall content of this disclosure.

[0034] Additionally, terms such as "first," "second," etc., may be used to describe various components, but the components should not be limited by these terms. These terms are used for the purpose of distinguishing one component from another.

[0035] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are "directly connected," but also cases where they are "electrically connected" with other components interposed between them. Furthermore, when a part is described as "including" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0036] Phrases such as "in one embodiment" appearing in various places in this disclosure do not necessarily refer to the same embodiment.

[0037] One embodiment of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various numbers of hardware and / or software configurations that execute specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a specific function. Additionally, for example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages. The functional blocks may be implemented as algorithms executed on one or more processors. Furthermore, the present disclosure may employ prior art for electronic configuration, signal processing, and / or data processing, etc. Terms such as "mechanism," "element," "means," and "configuration" may be used broadly and are not limited to mechanical and physical configurations.

[0038] Furthermore, the connecting lines or connecting members between the components depicted in the drawings are merely illustrative of functional connections and / or physical or circuit connections. In the actual device, connections between components may be represented by various alternative or added functional connections, physical connections, or circuit connections.

[0039] In this document, the preview image may be an image acquired through a camera of an electronic device and displayed in real time on the display of the electronic device.

[0040] Additionally, the text within the preview image may be the text to be translated, and the translated text may be the text containing the result of translating the text.

[0041] Additionally, the visual object may be a visual element that includes translated text and is overlaid on the preview image. The visual object may be created, for example, in the form of a window containing the translated text, but is not limited thereto. The visual object may be created, for example, by considering the shape and color of an object containing text within the preview image. Additionally, the visual object may include, for example, additional information retrieved in relation to the text.

[0042] The present disclosure will be described in detail below with reference to the attached drawings.

[0043] FIG. 1 is a diagram showing an overview of an electronic device according to one embodiment providing a translation result for text on an image.

[0044] Referring to FIG. 1, an electronic device (1000) according to one embodiment may generate a visual object (16) containing a translation result of text within a preview image (10) and overlay it on the preview image (10). When text disappears within the preview image, the electronic device (1000) according to one embodiment may maintain, change, or remove the display of the visual object (16) according to a preset condition.

[0045] Referring to identification number 1-1, an electronic device (1000) according to one embodiment may provide a translation result for text within a preview image (10). The electronic device (1000) displays a preview image (10) including a sign (12) and a person (14), and the sign (12) may include "STOP" which is the text to be translated.

[0046] Referring to identification number 1-2, according to one embodiment, an electronic device (1000) can obtain a translated text "stop" by translating the text "STOP" in a preview image (10), create a visual object (16) containing the translated text, and display the visual object (16) on the preview image (10).

[0047] Referring to identification numbers 1-3 and 1-4, the electronic device (1000) determines whether the sign (12) in the preview image (10) has disappeared due to a preset condition, and based on the result of the determination, can change or remove the display of the visual object (12).

[0048] The electronic device (1000) can identify when a person (14) moves toward a sign (12) and the sign (12) is obscured by the person (14), and the electronic device (1000) can change the color of a visual object (916) based on the color of the person (14).

[0049] Referring to identification numbers 1-4, the electronic device (1000) can identify that the shooting direction of the camera is changed by the movement of the electronic device (1000). Additionally, as the shooting direction of the camera is changed, the electronic device (1000) can remove the visual object (12) from the preview image (910).

[0050] According to one embodiment, the electronic device (1000) may be a smartphone, tablet PC, PC, smart TV, mobile phone, PDA (personal digital assistant), laptop, media player, micro server, GPS (global positioning system) device, e-book terminal, digital broadcasting terminal, navigation, kiosk, MP3 player, digital camera, home appliance, and other mobile or non-mobile computing devices, but is not limited thereto. Additionally, the electronic device (1000) may be a wearable device such as a watch, glasses, hair band, and ring equipped with communication functions and data processing functions. However, it is not limited thereto, and the electronic device (1000) may include any type of device capable of displaying a preview image.

[0051] FIG. 2 is a flowchart of a method in which an electronic device according to one embodiment provides a translation result on an image.

[0052] In operation 210, an electronic device (1000) according to one embodiment can acquire a preview image through a camera. The electronic device (1000) can activate the camera of the electronic device (1000) by running a camera application and can acquire a preview image through the activated camera. The electronic device (1000) can acquire a preview image in real time through the camera and display it on the screen of the electronic device (1000).

[0053] In operation 220, an electronic device (1000) according to one embodiment may obtain a first translated text by translating a first text within a preview image. The first text may be text that is the subject of translation, for example, text displayed on a specific object within the preview image. Additionally, the first translated text may be a translation result obtained by translating the first text. For example, the electronic device (1000) may identify the first text from the preview image using an optical character recognition (OCR) function and input the identified first text into an artificial intelligence model trained for translation, but is not limited thereto. In this case, the artificial intelligence model for translation may be an artificial intelligence model that performs translation based on text. Alternatively, for example, the electronic device (1000) may input the preview image into an artificial intelligence model trained for translation and obtain a translation result output from the artificial intelligence model. In this case, the artificial intelligence model for translation may be an artificial intelligence model that recognizes and translates text within an image.

[0054] In operation 230, an electronic device (1000) according to one embodiment may generate a first visual object including a first translation text based on a preview image, a first text, and a first translation text. According to one embodiment, the first visual object may be an object generated to include the first translation text and displayed on the screen of the electronic device (1000). For example, the first visual object may be generated to include the first translation text and a background. For example, the first visual object may be generated to include an image within the preview image, the first translation text, and a background. For example, the first visual object may be generated in the form of a window that is displayed separately from the preview image. For example, the first visual object may include a first translation text having a determined size and color, and may have a determined shape, size, and background color. For example, the background color of the first visual object may have the same or similar color as the color of the first object containing the first text. For example, the color of the first translation text within the first visual object may have a color distinct from the background color.

[0055] According to one embodiment, the electronic device (1000) may apply a preview image together with at least one of a first text or a first translation text to a first artificial intelligence model trained for generating a visual object. The first artificial intelligence model may be, for example, an artificial intelligence model trained to generate first visual objects having colors that are easy to visually recognize by a user, by considering the color of the first text included in the preview image, the color of the first object on which the first text is displayed, and the color of other objects within the preview image. For example, the first artificial intelligence model may be a multimodal model that learns and processes relationships between data of various types or various modalities, such as text and images. The first artificial intelligence model may be a large multimodal model (LMM) trained using text data and image data. The first artificial intelligence model may generate a visual object containing a translation text based on a text prompt, an image prompt, or a prompt composed of text and an image input to the first artificial intelligence model.

[0056] According to one embodiment, an electronic device (1000) may acquire a first visual object generated by taking into account the user's personal information. The user's personal information may include the user's profile information and / or information regarding the user's visual characteristics. Additionally, the user's visual characteristics information may include, for example, at least one of information regarding the user's eyesight, information regarding whether the user is colorblind, or information regarding the user's visual preferences. Additionally, the information regarding the user's visual preferences may include, for example, information regarding the color preferred by the user, the font preferred by the user, the text size preferred by the user, the font set on the user's electronic device (1000) and / or another electronic device, and the text size set on the user's electronic device (1000) and / or another electronic device, but is not limited thereto.

[0057] In this case, the electronic device (1000) can obtain a first visual object generated by taking into account the user's personal information by applying the user's personal information and a preview image together with at least one of a first text or a first translation text to a first artificial intelligence model. In this case, the first artificial intelligence model may be an artificial intelligence model trained to generate first visual objects having colors that are easy to visually recognize by the user, taking into account, for example, the color of the first text included in the preview image, the color of the first object on which the first text is displayed, the color of other objects within the preview image, and the user's visual characteristics.

[0058] An example of an electronic device (1000) generating a first visual object using a first artificial intelligence model will be described in more detail in FIGS. 3 and 4.

[0059] In operation 240, an electronic device (1000) according to one embodiment may display a first visual object by overlapping it on a preview image. The electronic device (1000) may display the first visual object on the first text so as to overlap at least partially with the first text. The electronic device (1000) may display the first visual object so as to overlap at least partially with a first object containing the first text.

[0060] According to one embodiment, the electronic device (1000) can display a first visual object so that the first visual object is superimposed on the first text in real time by monitoring the position of the first text. In this case, the electronic device (1000) can identify the position of the first text within the preview image in real time and can move the display position of the first visual object in real time as the first text moves within the preview image.

[0061] According to one embodiment, the electronic device (1000) can render a preview image and a first visual object so that the preview image and the first visual object are superimposed and displayed.

[0062] In operation 250, an electronic device (1000) according to one embodiment can identify whether at least a portion of the first text has disappeared from the preview image according to a preset condition. According to one embodiment, the electronic device (1000) can identify whether at least a portion of the first text has disappeared from the preview image according to the user's intention.

[0063] According to one embodiment, the electronic device (1000) can identify whether at least a portion of the first text has disappeared from the preview image according to the active intent of the user. For example, the electronic device (1000) can identify whether at least a portion of the first text has disappeared from the preview image due to the movement of the electronic device (1000). For example, the user can move the electronic device (1000) to change the shooting direction of the electronic device (1000), and as the direction in which the camera of the electronic device (1000) is facing changes, at least a portion of the first text may disappear within the preview image. In this case, the electronic device (1000) can identify whether at least a portion of the first text has disappeared from the preview image due to the movement of the electronic device (1000), for example, based on at least one of the distance moved by the electronic device (1000), the direction of movement, and the degree to which the shooting direction of the electronic device (1000) has changed.

[0064] According to one embodiment, the electronic device (1000) can identify whether at least a portion of the first text has disappeared from the preview image due to the user's passive intention. For example, the electronic device (1000) can identify whether at least a portion of the first text has disappeared from the preview image due to the movement of the first object. For example, when the first object, which is a subject in the real world, moves, the user may not change the shooting direction of the electronic device (1000). Accordingly, the position of the first object within the preview image may move, and the first object may disappear within the preview image. In this case, the electronic device (1000) can identify whether at least a portion of the first text has disappeared from the preview image due to the movement of the first object by, for example, monitoring the relative position of the first object and another object within the preview image.

[0065] According to one embodiment, the electronic device (1000) can identify whether at least a portion of the first text has disappeared in the preview image according to a preset condition by comparing the degree of movement of the electronic device (1000) with a first threshold and the degree of movement of the first object with a second threshold. For example, if the degree of movement of the electronic device (1000) (e.g., distance moved, direction of movement, amount of rotation, direction of rotation, etc.) is greater than the first threshold, or if the degree of movement of the first object (e.g., distance moved, direction of movement within the preview image, etc.) is greater than the second threshold, the electronic device (1000) can identify whether at least a portion of the first text has disappeared in the preview image according to a preset condition, but is not limited thereto.

[0066] According to one embodiment, if at least a portion of the first text disappears from the preview image according to a preset condition, the electronic device (1000) in operation 260 may remove the first visual object on the preview image. If at least a portion of the first text disappears from the preview image according to a preset condition, the electronic device (1000) may determine that the first text has disappeared from the preview image according to the user's intention and may not display the first translation text on the preview image. For example, the first text may disappear from the preview image by the user changing the direction of the camera of the electronic device (1000) to a different direction, in which case the electronic device (1000) may stop rendering the first visual object containing the first translation text for the first text. For example, the first text may disappear from the preview image by the first object containing the first text moving on its own, in which case the electronic device (1000) may stop rendering the first visual object containing the first translation text for the first text.

[0067] According to one embodiment, in the case where at least a portion of the first text has not disappeared from the preview image according to a preset condition, the electronic device (1000) in operation 270 may maintain the first visual object on the preview image. For example, the first object containing the first text may be obscured by another object, and in this case, it may be determined that the first text is not being displayed in the preview image regardless of the user's intention. In this case, even while the first text is not being displayed in the preview image, the electronic device (1000) may maintain the display of the first visual object containing the first translated text.

[0068] FIG. 3a is a diagram illustrating an example in which a visual object based on a preview image, a first text, and a first translation text is obtained through an artificial intelligence model according to one embodiment.

[0069] Referring to FIG. 3a, according to one embodiment, an input prompt including a preview image, a first text, and a first translation text may be input to a first artificial intelligence model (1090). In this case, for example, the input prompt may include a command requesting the creation of a first visual object including the first translation text, and the first translation text, including the preview image, the first text, and the first translation text. The input prompt may include, for example, text commanding the creation of a first visual object including the first translation text so that the first translation text, which is a translation of the first text in the preview image, is displayed distinctly on the preview image, but is not limited thereto.

[0070] Additionally, for example, the input prompt may be generated in a text format, but is not limited thereto. For example, the input prompt may be generated as data embedded in a preset format that can be input into the first artificial intelligence model. In this case, the preview image, the first text, the first translation text, and the command may be embedded so that they can be input into the first artificial intelligence model (1090) and input into the first artificial intelligence model (1090).

[0071] For example, a first artificial intelligence model that receives an input prompt can generate a first visual object by considering the colors of objects within a preview image so that the first translation text can be distinguished.

[0072] Meanwhile, according to one embodiment, an input prompt including a preview image and a first text may be input to the first artificial intelligence model (1090). In this case, for example, the input prompt may include a command requesting the creation of a first visual object that includes a preview image and a first text and includes a translation result of the first text.

[0073] The input prompt may include, for example, text that commands the creation of a first visual object containing the first translation text so that the first translation text, which is a translation of the first text in the preview image, is displayed distinctly on the preview image, but is not limited thereto.

[0074] Additionally, for example, the input prompt may be generated in a text format, but is not limited thereto. For example, the input prompt may be generated as data embedded in a preset format that can be input into the first artificial intelligence model. In this case, the preview image, the first text, and the command may be embedded so that they can be input into the first artificial intelligence model (1090) and input into the first artificial intelligence model (1090).

[0075] In addition, for example, a first artificial intelligence model that receives an input prompt can translate a first text and generate a first visual object that can distinguish the first translated text by considering the colors of objects within a preview image.

[0076] Meanwhile, according to one embodiment, an input prompt including a preview image may be input to the first artificial intelligence model (1090). In this case, for example, the input prompt may include a command requesting the creation of a first visual object that includes a preview image and a translation result of text within the preview image.

[0077] The input prompt may include, for example, text that commands the creation of a first visual object so that the translation result of translating text within the preview image is displayed distinctly on the preview image, but is not limited thereto.

[0078] Additionally, for example, the input prompt may be generated in a text format, but is not limited thereto. For example, the input prompt may be generated as data embedded in a preset format that can be input into the first artificial intelligence model. In this case, the preview image and command may be embedded so that they can be input into the first artificial intelligence model (1090) and input into the first artificial intelligence model (1090).

[0079] In addition, for example, a first artificial intelligence model that receives an input prompt can identify text within a preview image, translate the identified text, and generate a first visual object that allows the translated text to be distinguished by considering the colors of objects within the preview image.

[0080] FIG. 3b is a diagram illustrating an example in which a visual object is obtained through an artificial intelligence model by reflecting the user's personal information according to one embodiment.

[0081] Referring to FIG. 3b, according to one embodiment, an input prompt including a preview image, a first text, a first translation text, and personal information of a user may be input into a first artificial intelligence model (1090). The personal information of a user may include, for example, information regarding the user's visual characteristics. Additionally, the information regarding the user's visual characteristics may include, for example, at least one of information regarding the user's eyesight, information regarding whether the user is colorblind, or information regarding the user's visual preferences, but is not limited thereto. Additionally, the information regarding the user's visual preferences may include, for example, information regarding the color preferred by the user, the font preferred by the user, the text size preferred by the user, the font set on the user's electronic device (1000) and / or other electronic devices, and the text size set on the user's electronic device (1000) and / or other electronic devices, but is not limited thereto. Additionally, for example, the personal information of a user may include profile information such as the user's gender, age, and occupation.

[0082] In this case, for example, the input prompt may include a preview image, a first text, a first translation text, and the user's personal information, and may include a command requesting the creation of a first visual object containing the first translation text. The input prompt may include, for example, text instructing the creation of a first visual object containing the first translation text so that the first translation text, which is a translation of the first text in the preview image, is displayed distinctly on the preview image, but is not limited thereto.

[0083] Additionally, for example, the input prompt may be generated in a text format, but is not limited thereto. For example, the input prompt may be generated as data embedded in a preset format that can be input into the first artificial intelligence model. In this case, the preview image, the first text, the first translation text, personal information, and the command may be embedded so that they can be input into the first artificial intelligence model (1090) and input into the first artificial intelligence model (1090).

[0084] Additionally, for example, a first artificial intelligence model that receives an input prompt can generate a first visual object so that the first translation text can be easily distinguished by the user by taking into account the user's eyesight, whether the user is colorblind, the user's visual preferences, and / or the colors of objects within the preview image.

[0085] FIG. 3c is a diagram illustrating an example in which a visual object including a first translation text and additional information is obtained through an artificial intelligence model according to one embodiment.

[0086] Referring to FIG. 3c, according to one embodiment, an input prompt including a preview image, a first text, and a first translation text may be input to a first artificial intelligence model (1090). In this case, for example, the input prompt may include a command requesting the creation of a first visual object that includes a preview image, a first text, and a first translation text, and includes additional information related to the first translation text and the first text. The input prompt may include, for example, text that commands the creation of a first visual object including the first translation text so that the first translation text translated from the first text in the preview image is displayed distinctly on the preview image, and additional information related to the first text within the first visual object, but is not limited thereto.

[0087] Additionally, for example, the input prompt may be generated in a text format, but is not limited thereto. For example, the input prompt may be generated as data embedded in a preset format that can be input into the first artificial intelligence model. In this case, the preview image, the first text, the first translation text, and the command may be embedded so that they can be input into the first artificial intelligence model (1090) and input into the first artificial intelligence model (1090).

[0088] Additionally, for example, a first artificial intelligence model that receives an input prompt may generate a first visual object by considering the colors of objects within a preview image so that the first translation text and additional information can be distinguished. The additional information may include, for example, information retrieved in relation to the first text and information for additionally explaining the first text. Additionally, the additional information may include, for example, at least one of text corresponding to the content of the additional information, an image, or link information that allows access to the additional information, but is not limited thereto.

[0089] FIG. 4a is a diagram showing an example in which a visual object generated from an artificial intelligence model according to one embodiment is displayed on a preview image.

[0090] Referring to FIG. 4a, a preview image (40), a first text "STOP" within the preview image (40), and a first translation text "stop" for the first text can be applied to a first artificial intelligence model (1090). For example, a preview image (40), a first text "STOP" within the preview image (40), and a first translation text "stop" for the first text can be input into the first artificial intelligence model (1090). Alternatively, a preview image (40), a first text "STOP" within the preview image (40), and a first translation text "stop" for the first text can be embedded, and the embedded data can be input into the first artificial intelligence model (1090).

[0091] Subsequently, the first visual object (42) can be output from the first artificial intelligence model (1090). The color and size of the first visual object (42) can be determined by considering the colors of the objects in the preview image (42) and the color of the first text. For example, the first visual object (42) can be created to have a red background color and a white text color by considering the color of the first text "STOP" and the color around the first text. Additionally, for example, the size and shape of the first visual object (42) can be determined by considering the size of the first object containing the first text.

[0092] Subsequently, the electronic device (1000) can display the first visual object (42) superimposed on the preview image (40). The electronic device (1000) can render the preview image (40) and the first visual object (42) so that the first visual object (42) is overlaid on the preview image (40). Accordingly, the first visual object (42) can be displayed superimposed on the first text "STOP".

[0093] FIG. 4b is a diagram showing an example in which a visual object generated from an artificial intelligence model based on user visual characteristic information according to one embodiment is displayed on a preview image.

[0094] Referring to FIG. 4b, a preview image (40), a first text "STOP" within the preview image (40), a first translation text "Stop" for the first text, information regarding the user's eyesight, and information regarding the user's color blindness may be applied to a first artificial intelligence model (1090). For example, a preview image (40), a first text "STOP" within the preview image (40), a first translation text "Stop" for the first text, information regarding the user's eyesight, and information regarding the user's color blindness may be input into the first artificial intelligence model (1090). Alternatively, a preview image (40), a first text "STOP" within the preview image (40), a first translation text "Stop" for the first text, information regarding the user's eyesight, and information regarding the user's color blindness may be embedded, and the embedded data may be input into the first artificial intelligence model (1090). In this case, for example, information regarding the user's eyesight may have a value indicating that the user has poor farsightedness, and information regarding the user's color blindness may have a value indicating that the user has red-green color blindness.

[0095] Subsequently, the first visual object (43) may be output from the first artificial intelligence model (1090). The color and size of the first visual object (43) may be determined by considering the colors of the objects in the preview image (43), the color of the first text, the user's eyesight, and whether the user is colorblind. For example, the first visual object (43) may be created to have a black background color and a white text color by considering the color of the first text "STOP", the color around the first text, the user's raw eyesight, and the user's red-green colorblindness. Additionally, for example, the size of the first visual object (43) may be determined to have a size larger than the size of the first object containing the first text by considering the user's raw eyesight.

[0096] Subsequently, the electronic device (1000) can display the first visual object (43) superimposed on the preview image (40). The electronic device (1000) can render the preview image (40) and the first visual object (43) so that the first visual object (43) is overlaid on the preview image (40). Accordingly, the first visual object (43) can be displayed superimposed on the first text "STOP".

[0097] In FIG. 4b, it is described that the first visual object (43) is determined based on information regarding visual acuity and color blindness as visual characteristic information of the user, but is not limited thereto. For example, the first visual object (43) may be determined based on personal information of the user, such as information regarding the user's visual preference and user profile information.

[0098] FIG. 4c is a drawing showing an example of a visual object including a translation result for text in a preview image and images in a preview image according to one embodiment.

[0099] Referring to identification number 4c-1 in FIG. 4c, the preview image (45) may include texts (46-1a, 46-2a, 46-3a) and a plurality of images (46-4, 46-5) to be translated. For example, the texts (46-1a, 46-2a, 46-3a) may include text describing a method of assembling chair legs. Additionally, for example, the plurality of images (46-4, 46-5) may include drawings illustrating the process of assembling chair legs.

[0100] Referring to identification number 4c-2 in FIG. 4c, a visual object (46) may be displayed superimposed on a preview image (45). The visual object (46) may include translated texts (46-1b, 46-2b, 46-3b) generated by translating texts (46-1a, 46-2a, 46-3a). Additionally, the visual object (46) may include a plurality of images (46-4, 46-5). The translated texts (46-1b, 46-2b, 46-3b) and the plurality of images (46-4, 46-5) may be placed within the visual object (46) to correspond to the composition of the preview image.

[0101] Additionally, for example, the visual object (46) may include additional information (46-7) related to the texts (46-1a, 46-2a, 46-3a). For example, the additional information (46-7) may be displayed in the form of an object (e.g., button, icon, etc.) for accessing a video of assembling a chair. For example, the additional information (46-7) may include link information for downloading a video of assembling a chair.

[0102] FIG. 4d is a drawing showing an example of a visual object including translation results and additional information for text within a preview image according to one embodiment.

[0103] Referring to identification number 4d-1 in FIG. 4d, the preview image (47) may include an object (48-1) containing text to be translated. For example, the object (48-1) may be an open book.

[0104] Referring to identification number 4d-2 in FIG. 4d, a visual object (48-2) can be displayed superimposed on a preview image (47). The visual object (48-2) may include translated text (48-3) generated by translating text within an object (48-1).

[0105] Additionally, the visual object (48-2) may include additional information (48-4) that is additionally retrieved based on the text within the object (48-1). For example, if the text within the object (48-1) is related to Dabotap, the additional information (48-4) included in the visual object (48-2) may include an image of Dabotap.

[0106] FIG. 4e is a drawing showing an example in which a visual object is displayed when the electronic device according to one embodiment is a wearable electronic device.

[0107] Referring to FIG. 4e, an electronic device (1000) according to one embodiment may be a head-mounted device (HMD) device including a display or smart glasses. In this case, the electronic device (1000) may display an augmented reality image. For example, the electronic device (1000) may identify the location of an object relative to the electronic device (1000) and the location of the user's eyes relative to the electronic device (1000), and, taking into account the location of the object and the location of the eyes, display a visual object (49-1, 49-2) containing translated text on the screen of the electronic device (1000). Accordingly, the visual object (49-1, 49-2) may be shown to the user as if superimposed on a real-world object.

[0108] FIG. 5 is a flowchart of a method for determining whether an electronic device according to one embodiment maintains the display of a first visual object when the first text, which is the subject of translation, disappears from the preview image.

[0109] The operations of FIG. 5 may correspond to operations 250 to 270 of FIG. 2.

[0110] In operation 510, the electronic device (1000) can monitor the movement of the electronic device (1000). According to one embodiment, the electronic device (1000) can monitor the movement of the electronic device (1000) while a preview image is being displayed. For example, the electronic device (1000) can monitor the movement of the electronic device (1000) using a sensor of the electronic device (1000) as a first visual object is displayed on the preview image. For example, the electronic device (1000) can sense the movement of the electronic device (1000) using a motion sensor and / or a gyroscope sensor. Additionally, for example, the electronic device (1000) can monitor a change in the direction in which the camera of the electronic device (1000) is facing based on the sensed value.

[0111] In operation 520, the electronic device (1000) can monitor the movement of a first object containing a first text. According to one embodiment, the electronic device (1000) can monitor whether the first object containing the first text has moved within the preview image while the preview image is being displayed. For example, the electronic device (1000) can monitor whether the first object has moved within the preview image by analyzing the preview image as the first visual object is displayed on the preview image.

[0112] According to one embodiment, the electronic device (1000) can apply a preview image in real time to an artificial intelligence model trained for object recognition, and can monitor the positions of the first object and other objects within the preview image based on the first object and other objects recognized using the artificial intelligence model. Additionally, the electronic device (1000) can determine whether the first object, which is the subject, has moved based on the relationship. In this case, the other objects compared with the first object may be non-moving subjects (e.g., trees, buildings, etc.).

[0113] In operation 530, the electronic device (1000) can determine whether the first text has disappeared from the preview image due to the movement of the electronic device (1000). According to one embodiment, the electronic device (1000) can determine whether the first text has disappeared from the preview image due to the movement of the electronic device (1000) when the first text has disappeared from the preview image.

[0114] According to one embodiment, the electronic device (1000) can determine whether the first text has disappeared from the preview image based on the amount of movement of the electronic device (1000). For example, the electronic device (1000) can identify the degree to which the direction of the camera has changed due to the movement of the electronic device (1000). Additionally, for example, the electronic device (1000) can calculate the amount of change in the position of the first text within the preview image based on the degree to which the direction of the camera has changed, the magnification of the camera, and the angle of view of the camera. Additionally, for example, if the amount of change in the position of the first text is greater than the distance from the position of the first text to the edge of the preview image, the electronic device (1000) can determine that the first text has disappeared from the preview image due to the movement of the electronic device (1000).

[0115] According to one embodiment, if it is determined that the first text has disappeared from the preview image due to the movement of the electronic device (1000), the electronic device (1000) may perform operation 260. Additionally, if it is determined that the first text has not disappeared from the preview image due to the movement of the electronic device (1000), the electronic device (1000) may perform operation 540.

[0116] In operation 540, the electronic device (1000) can determine whether the first text has disappeared from the preview image due to the movement of the first object containing the first text. According to one embodiment, the electronic device (1000) can determine whether the first text has disappeared from the preview image due to the movement of the first object containing the first text when the first text has disappeared from the preview image.

[0117] According to one embodiment, the electronic device (1000) can determine to what extent the position of the first object within the preview image has changed due to the movement of the first object, which is the subject, based on the positional relationship between the first object and another fixed object within the preview image. Additionally, if the amount of change in the position of the first object within the preview image due to the movement of the first object is greater than the distance from the position of the first text to the border of the preview image, the electronic device (1000) can determine that the first text has disappeared from the preview image due to the movement of the first object.

[0118] According to one embodiment, if it is determined that the first text has not disappeared from the preview image due to the movement of the first object, the electronic device (1000) can perform operation 270.

[0119] Additionally, according to one embodiment, if it is determined that the first text has disappeared from the preview image due to the movement of the first object, the electronic device (1000) can perform operation 550.

[0120] In operation 550, the electronic device (1000) can determine whether the first text disappears from the preview image by another object as the first object moves. According to one embodiment, as the first object moves toward another object located closer to the electronic device (1000) than the first object, at least a portion of the first object may be obscured by the other object. For example, at least a portion of the first object may be obscured by another object while the first object is moving. In this case, the other object may be a stationary object or a moving object. Additionally, as at least a portion of the first object is obscured by another object, at least a portion of the first text may be obscured by the other object.

[0121] According to one embodiment, if it is determined that the first text has disappeared from the preview image by another object as the first object moves, the electronic device (1000) may perform operation 270. Additionally, if it is determined that the first text has not disappeared from the preview image by another object as the first object moves, the electronic device (1000) may perform operation 260.

[0122] Although it has been described above that operation 540 is performed after operation 530, it is not limited thereto. For example, operations 530 and 540 may be performed together, in which case it may be determined whether the first text has disappeared from the preview image based on whether the electronic device (1000) has moved and / or whether the first object has moved. Additionally, according to one embodiment, operations 530 to 550 may be performed when at least a portion of the first text has disappeared from the preview image.

[0123] According to one embodiment, for example, at least a portion of the first text disappearing from the preview image may include at least a portion of the first text disappearing from the preview image by moving outside the field of view range of the camera of the electronic device (1000). Additionally, for example, at least a portion of the first text disappearing from the preview image may include at least a portion of the first text being obscured by another object by at least a portion of the object containing the first text being obscured by another object within the preview image.

[0124] According to one embodiment, for example, if the first object disappears from the preview image because the first object moves outside the field of view range of the camera of the electronic device (1000), the first visual object may be removed from the preview image. Additionally, for example, if the first object is obscured by another object within the preview image because the first object moves within the field of view range of the camera of the electronic device (1000), the display of the first visual object may be maintained. Additionally, for example, if the first object disappears from the preview image because the first object is positioned outside the field of view range of the camera of the electronic device (1000) due to the movement of the electronic device (1000), the first visual object may be removed from the preview image. Additionally, for example, if another object obscures the first object within the preview image because another object moves, the display of the first visual object may be maintained.

[0125] FIG. 6a is a drawing showing an example in which a first visual object is displayed when the first text is obscured by another object as the first object containing the first text moves according to one embodiment.

[0126] Referring to Fig. 6a, a scene of a car with the word "victory" engraved on its door driving on a road can be captured by an electronic device (1000).

[0127] Referring to identification number 6a-1, the preview image (60) may include a first object (62), which is a car driving on a road, and a first text (64-1) of "victory" engraved on the door of the car. Additionally, the first object (62) may be located in the left area within the preview image (60).

[0128] Referring to identification number 6a-2, a first visual object (64-2) containing the first translation text "victory" may be displayed overlaid on the first text (64-1). Additionally, as the vehicle travels on the road, the position of the first object (62) and the position of the first visual object (64-2) may be moved within the preview image (60). For example, the first object (62) and the first visual object (64-2) may move to the right area within the preview image (60). In this case, the electronic device (1000) may anticipate that the first object (62) will move and be obscured by another object (69) and may decide to maintain the display of the first visual object (64-2). For example, the other object (69) may be a tree located closer to the electronic device (1000) than the first object (62).

[0129] Additionally, even if the first object (62) is obscured by another object (69) due to the movement of the first object (62), the electronic device (1000) can maintain the display of the first visual object (64-2). Accordingly, the first visual object (64-2) can be overlapped on the other object (69). For example, as the first object (62) containing the first text (64-1) moves toward the other object (69), the first text (64-1) may disappear from the preview image as it is obscured by the other object (69). In this case, the electronic device (1000) can overlap the first visual object (64-2) on the first object (62) and the other object (69).

[0130] Referring to identification number 6a-3, as the vehicle moves on the road, the vehicle may move out of the field of view of the camera of the electronic device (1000). Accordingly, the first object (62) may disappear from the preview image (60). The electronic device (1000) identifies that the first object (62) has disappeared from the preview image (60) due to the movement of the vehicle and can remove the first visual object (64-2) from the preview image (60).

[0131] FIG. 6b is a drawing showing an example of stopping the display of a first visual object when the first text disappears from the preview image as the shooting direction of the electronic device changes.

[0132] Referring to FIG. 6b, a scene including a sign (67) with the first text "STOP" engraved on it can be captured by an electronic device (1000).

[0133] Referring to identification number 6b-1, the preview image (66) may include a sign (67) containing the first text "STOP".

[0134] Referring to identification number 6b-2, a first visual object (68) containing the first translation text "stop" for the first text "STOP" can be overlaid on a preview image (66). The first visual object (68) can be displayed on the first text within the sign (67). Additionally, for example, the first visual object (68) may have a background of the same color as the sign (67).

[0135] Referring to identification number 6b-3, as the shooting direction of the electronic device (1000) changes, the sign (67) containing the first text may disappear from the preview image (66). For example, the electronic device (1000) can identify that the shooting direction of the electronic device (1000) has changed by monitoring the movement of the electronic device (1000). Additionally, the electronic device (1000) can determine that the first text of the sign (67) has disappeared from the preview image (66) due to the movement of the electronic device (1000) based on the degree of change in the shooting direction of the electronic device (1000). The electronic device (1000) can identify that the first text of the sign (67) has disappeared from the preview image (66) due to the movement of the electronic device (1000) and remove the first visual object (68) from the preview image (66).

[0136] FIG. 7 is a flowchart of a method in which an electronic device according to one embodiment predicts whether a first object containing a first text disappears due to another object and performs a display setting for a first visual object.

[0137] The operations of FIG. 7 may be performed after operation 240 of FIG. 2, but are not limited thereto.

[0138] In operation 710, the electronic device (1000) can predict whether the first object will be obscured by the other object as the other object moves. According to one embodiment, the electronic device (1000) can monitor the movement and depth changes of the first object and the other object by recognizing the first object and the other object in real time within a preview image. The electronic device (1000) can apply the preview image in real time to an artificial intelligence model trained to recognize objects and the depth of objects within an image, for example. Additionally, for example, the electronic device (1000) may recognize the depth of an object using a sensor (e.g., a TOF sensor) to detect the depth of an object. Additionally, the electronic device (1000) can identify changes in position and depth of the first object and the other object within the preview image based on output data from the artificial intelligence model.

[0139] Additionally, the electronic device (1000) can predict whether the first object will be obscured by another object based on changes in position and depth of the first object and other objects. For example, the electronic device (1000) can identify that another object with a shallower depth than the first object is moving toward the first object within a preview image, and can predict whether the first object will be obscured by the other object based on the movement path and speed of the other object.

[0140] In operation 720, the electronic device (1000) can set the first visual object as the top layer. According to one embodiment, if it is predicted that the first object will be obscured by another object within the preview image, the electronic device (1000) can adjust the overlap level of the first visual object. For example, the electronic device (1000) can set the first visual object to be displayed as the top layer. For example, as the first visual object is set as the top layer, the first visual object can be displayed above other visual objects.

[0141] FIG. 8 is a drawing showing an example in which an electronic device according to one embodiment changes and displays the attributes of a first visual object.

[0142] Referring to FIG. 8, a scene including a sign (81) with the first text "STOP" engraved on it and a person (84) can be captured by an electronic device (1000).

[0143] Referring to identification numbers 8-1 and 8-2, for example, the preview image (80) may include a sign (81) containing the first text "STOP" and a person (84). Additionally, within the preview image (84), the person (84) may move toward the sign (81).

[0144] Additionally, a first visual object (82) containing "stop," which is a translation text for "STOP," may be overlaid on the preview image (80). The first visual object (82) may be displayed on the first text of the sign (81). Additionally, the first visual object (82) may have a red background that is the same color as the sign (81).

[0145] Referring to identification number 8-3, as a person (84) moves toward the sign (81), the first text of the sign (81) may be obscured by the person (84). As the first text of the sign (81) is obscured by the person (84), the first text may disappear from the preview image (80).

[0146] Additionally, the electronic device (1000) can identify that the first text has disappeared due to another object. Additionally, it can identify that the first text has not disappeared due to the movement of the electronic device (1000) and / or the movement of the sign (81).

[0147] In this case, the electronic device (1000) may decide to maintain the display of the first visual object (82). Additionally, the electronic device (1000) may change the color of the first visual object (82) based on the color of another object, a person (84). For example, the first visual object (82) may change the background color of the first visual object (82) and the text color of the first translation text based on the color of the person (84).

[0148] According to one embodiment, the electronic device (1000) can predict whether a sign (81) containing the first text will be obscured by the person (84) by monitoring the movement of the person (84) within the preview image (80). When it is predicted that the sign (81) containing the first text will be obscured by the person (84), the electronic device (1000) can change the color of the first visual object (82) in advance.

[0149] FIG. 9 is a flowchart of a method for displaying visual objects when a plurality of objects within a preview image overlap, according to one embodiment of an electronic device.

[0150] In operation 900, the electronic device (1000) can obtain a second translated text by translating a second text within a preview image. The second text may be a text that is the subject of translation and may be a text different from the first text. For example, the second text may be a text displayed on a second object that is different from the first object in which the first text is displayed within the preview image. Additionally, the second translated text may be a translation result obtained by translating the second text. For example, the electronic device (1000) may use an optical character recognition (OCR) function to identify the second text from the preview image and input the text into an artificial intelligence model trained for translation, but is not limited thereto. In this case, the artificial intelligence model for translation may be an artificial intelligence model that performs translation based on text. Or, for example, the electronic device (1000) may input the preview image into an artificial intelligence model trained for translation and obtain a translation result output from the artificial intelligence model. In this case, the artificial intelligence model for translation may be an artificial intelligence model that recognizes and translates text within an image.

[0151] In operation 905, the electronic device (1000) may generate a second visual object containing the second translation text based on a preview image, a second text, and a second translation text. According to one embodiment, the second visual object may be an object generated to include the second translation text and displayed on the screen of the electronic device (1000). For example, the second visual object may be generated to include the second translation text and a background. For example, the second visual object may be generated to include an image within the preview image, the second translation text, and a background. For example, the second visual object may be generated in the form of a window that is displayed separately from the preview image. For example, the second visual object may include the second translation text having a determined size and color, and may have a determined shape, size, and background color. For example, the background color of the second visual object may have the same or similar color as the color of the second object containing the second text. For example, the color of the second translation text within the second visual object may have a color distinct from the background color.

[0152] According to one embodiment, the electronic device (1000) may apply a preview image together with at least one of a second text or a second translation text to a first artificial intelligence model trained for generating a visual object. The first artificial intelligence model may be, for example, an artificial intelligence model trained to generate second visual objects having colors that are easy to visually recognize by a user, taking into account the color of the second text included in the preview image, the color of the second object on which the second text is displayed, and the color of other objects within the preview image. For example, the first artificial intelligence model may be a multimodal model that learns and processes relationships between data of various types or various modalities, such as text and images. The first artificial intelligence model may be a large multimodal model (LMM) trained using text data and image data. The first artificial intelligence model may generate a visual object containing a translation text based on a text prompt, an image prompt, or a prompt composed of text and an image input to the first artificial intelligence model.

[0153] According to one embodiment, the electronic device (1000) may acquire a second visual object generated by taking into account the user's personal information. The user's personal information may include the user's profile information and / or information regarding the user's visual characteristics. Additionally, the user's visual characteristics information may include, for example, at least one of information regarding the user's eyesight, information regarding whether the user is colorblind, or information regarding the user's visual preferences. Additionally, the information regarding the user's visual preferences may include, for example, information regarding the color preferred by the user, the font preferred by the user, the text size preferred by the user, the font set on the user's electronic device (1000) and / or another electronic device, and the text size set on the user's electronic device (1000) and / or another electronic device, but is not limited thereto.

[0154] In this case, the electronic device (1000) can obtain a second visual object generated with consideration of the user's personal information by applying the user's personal information and a preview image together with at least one of a second text or a second translation text to a first artificial intelligence model. In this case, the first artificial intelligence model may be an artificial intelligence model trained to generate second visual objects having colors that are easy to visually recognize by the user, by considering, for example, the color of the second text included in the preview image, the color of the second object on which the second text is displayed, the color of other objects within the preview image, and the user's visual characteristics.

[0155] In operation 910, the electronic device (1000) can display a preview image and a second visual object by overlapping them.

[0156] The electronic device (1000) can display a second visual object on the second text so as to overlap at least partially with the second text. The electronic device (1000) can display a second visual object so as to overlap at least partially with a second object containing the second text.

[0157] According to one embodiment, the electronic device (1000) can display a second visual object so that the second visual object is superimposed on the second text in real time by monitoring the position of the second text. In this case, the electronic device (1000) can identify the position of the second text within the preview image in real time and can move the display position of the second visual object in real time as the second text moves within the preview image.

[0158] According to one embodiment, the electronic device (1000) can render a preview image and a second visual object so that the preview image and the second visual object are superimposed and displayed.

[0159] Although operations 900 to 910 have been described above as separate operations from operations 220 to 240 of FIG. 2, they are not limited thereto. If a second object containing a second text is added within the preview image, a first object containing a first text and a second object containing a second text may be included together within the preview image. In this case, operation 900 may be performed together with operation 220, operation 905 may be performed together with operation 230, and operation 910 may be performed together with operation 240. For example, a first translation text and a second translation text for the first text and the second text within the preview image may be obtained together. Additionally, for example, a first visual object and a second visual object may be obtained based on the preview image, the first text, the first translation text, the second text, and the second translation text. In addition, for example, a first visual object and a second visual object can be overlaid together on a preview image.

[0160] In operation 915, the electronic device (1000) can identify that a first object corresponding to a first text and a second object corresponding to a second text are superimposed. According to one embodiment, the first object and / or the second object may move within the preview image, and as the first object and / or the second object move, the first object and the second object may superimpose on each other. The electronic device (1000) can monitor the movement and depth changes of the first object and the second object by recognizing the first object and the second object within the preview image in real time. The electronic device (1000) can apply the preview image in real time to an artificial intelligence model trained to recognize objects and the depth of objects within the image, for example. Additionally, the electronic device (1000) can identify changes in position and depth of the first object and the second object within the preview image based on output data from the artificial intelligence model. Additionally, for example, the electronic device (1000) may recognize the depth of an object by using a sensor (e.g., a TOF sensor) to detect the depth of an object. Accordingly, the electronic device (1000) may determine that the first object and the second object overlap within the preview image. Additionally, the electronic device (1000) may predict that the first object and the second object overlap within the preview image.

[0161] In operation 920, the electronic device (1000) can identify whether the second translation text is associated with the first translation text. The electronic device (1000) can identify whether the first translation text and the second translation text are associated with each other based on whether the first text and the second text are related to each other. For example, if the first text and the second text are related to the same or similar subject, the electronic device (1000) can determine that the second translation text is associated with the first translation text. For example, if a word included in the first text is related to the same or similar category as a word included in the second text, the electronic device (1000) can determine that the second translation text is associated with the first translation text. For example, if the content of the first text and the content of the second text are related to the same or similar context, the electronic device (1000) can determine that the second translation text is associated with the first translation text. However, examples of the first translation text and the second translation text being associated are not limited thereto.

[0162] According to one embodiment, if it is determined in operation 920 that the second translation text is not associated with the first translation text, the electronic device (1000) may perform operation 925.

[0163] In operation 925, the electronic device (1000) can identify the nesting order of the first visual object and the second visual object.

[0164] According to one embodiment, the electronic device (1000) can identify the overlapping order of the first visual object and the second visual object in order to determine whether to overlay the second visual object on the first visual object or to overlay the first visual object on the second visual object.

[0165] According to one embodiment, the overlapping order of a first visual object and a second visual object may be determined according to specified criteria. For example, the overlapping order of the first visual object and the second visual object may be determined according to the depth of the first object and the depth of the second object. Alternatively, for example, when the first visual object is fixed as the top layer, the overlapping order may be determined such that the first visual object is displayed on the second visual object.

[0166] In operation 930, the electronic device (1000) may display a first visual object and a second visual object by overlapping them based on the identified overlapping order. According to one embodiment, the electronic device (1000) may overlay the second visual object on the first visual object or overlay the first visual object on the second visual object according to the identified overlapping order.

[0167] According to one embodiment, if it is determined in operation 920 that the second translation text is associated with the first translation text, the electronic device (1000) can perform operation 935.

[0168] In operation 935, the electronic device (1000) may generate a third visual object comprising a first translation text and a second translation text. According to one embodiment, the electronic device (1000) may generate a third visual object comprising the first translation text and the second translation text together. For example, the third visual object may be generated to include the first translation text, the second translation text, and a background. For example, the third visual object may be generated to include an image within a preview image, the first translation text, the second translation text, and a background. For example, the third visual object may be generated in the form of a window that is displayed separately from the preview image. For example, the third visual object may include the first translation text and the second translation text having a determined size and color, and may have a determined shape, size, and background color. For example, the background color of the third visual object may be determined based on the color of the first object containing the first text and the color of the second object containing the second text. According to one embodiment, the electronic device (1000) may apply a preview image to a first artificial intelligence model trained for generating a visual object, together with at least one of a first text, a second text, a first translation text, or a second translation text. The first artificial intelligence model may be, for example, an artificial intelligence model trained to generate third visual objects having colors that are easily visually recognized by a user, by considering the color of the first text included in the preview image, the color of the first object, the color of the second text, the color of the second object, and the color of other objects within the preview image. For example, the first artificial intelligence model may be a multimodal model that learns and processes relationships between data of various types or various modalities, such as text and images.The first artificial intelligence model may be a large multimodal model (LMM) trained using text data and image data. The first artificial intelligence model may generate a visual object containing translated text based on a text prompt, an image prompt, or a prompt composed of text and an image input to the first artificial intelligence model.

[0169] In operation 940, a third visual object may be displayed in place of the first visual object and the second visual object. The electronic device (1000) may remove the first visual object and the second visual object from the preview image and display the third visual object on the preview image. The electronic device (1000) may overlay the third visual object on the first object and the second object that overlap each other.

[0170] FIG. 10 is a drawing showing an example of visual objects being displayed when a plurality of objects within a preview image overlap according to one embodiment.

[0171] Referring to identification number 1001 and identification number 1002, according to one embodiment, a preview image (1010) may include a first object (1020) including a first text, and a first visual object (1022) including a first translation text for the first text may be overlaid on the first object (1020).

[0172] Referring to identification number 1003, a second object (1030) including a second text according to one embodiment may be additionally included in the preview image (1010), and the second object (1030) may be superimposed on the first object (1020).

[0173] Referring to identification number 1004, an electronic device (1000) according to one embodiment may generate a third visual object (1024) including a first translation text and a second translation text. Additionally, the first visual object (1022) may be removed from the preview image (1010), and the third visual object (1024) may be overlaid on the first object (1020) and the second object (1030).

[0174] FIG. 11 is a drawing showing an example of a visual object fixed to the top layer being moved according to one embodiment.

[0175] Referring to identification number 1101, a visual object (1110) overlaid on a preview image (1100) according to one embodiment may be fixed as the top layer. Additionally, while the visual object (1110) fixed as the top layer is displayed, a first object (1120-1) containing a first text and a second object (1120-2) containing a second text may be added within the preview image (1100). For example, as the shooting direction of the camera moves, the first object (1120-1) and the second object (1120-2) may be added within the preview image (1100).

[0176] Referring to identification number 1102, the electronic device (1000) may create a first visual object (1120-2) containing a first translation text for a first text within a first object (1120-1), and may create a second visual object (1121-2) containing a second translation text for a second text within a second object (1121-1). Additionally, the electronic device (1000) may overlay the first visual object (1120-2) on the first object (1120-1) and overlay the second visual object (1121-2) on the second object (1121-1). In this case, the electronic device (1000) can move the position of the visual object (1110) so that the visual object (1110) fixed at the top does not overlap with the first visual object (1120-2) and the second visual object (1121-2).

[0177] FIG. 12 is a drawing showing an example in which the transparency of a visual object is controlled according to one embodiment.

[0178] Referring to identification number 1201, a visual object (1210) overlaid on a preview image (1200) according to one embodiment may be fixed as the top layer. Additionally, while the visual object (1210) fixed as the top layer is displayed, a plurality of objects (1220) may be added within the preview image (1200). For example, a plurality of people gathered for a photo shoot may be added within the preview image (1200).

[0179] Referring to identification number 1202, as a plurality of objects (1220) within a preview image (1200) according to one embodiment overlap with a visual object (1210), the electronic device (1000) can adjust the transparency of the visual object (1210). For example, the electronic device (1000) can adjust the transparency of the visual object (1210) to a high level. In this case, for example, the electronic device (1000) can determine the importance of the plurality of objects (1220) within the preview image (1200) and determine the transparency of the visual object (1210) based on the importance.

[0180] FIG. 13 is a drawing showing an example of fixing a visual object to an upper layer according to one embodiment.

[0181] Referring to identification number 1301, a visual object (1330) set as a top layer according to one embodiment may be overlaid on a preview image (1310). When the visual object (1330) is set to be overlaid as a top layer, an icon (50) indicating that the visual object (1330) is set as a top layer may be included within the visual object (1330).

[0182] Referring to identification number 1301, a visual object (1330) with its top layer setting disabled can be overlaid on a preview image (1310). As the icon (50) is touched by the user, the top layer setting for the visual object (1330) can be disabled. Additionally, an icon (52) indicating that the top layer setting for the visual object (1330) has been disabled can be included within the visual object (1330). Additionally, as the icon (52) is touched by the user, the visual object (1330) can be set as the top layer.

[0183] According to one embodiment, a method for an electronic device (1000, 1401) to provide a translation result for text within a preview image comprises: an operation (210) of acquiring a preview image through a camera (1450); an operation (220) of acquiring a first translated text by translating a first text within the preview image; an operation (230) of generating a first visual object including the first translated text based on the preview image, the first text, and the first translated text; an operation (240) of displaying the preview image and the first visual object, wherein the first visual object is overlaid on the first text within the preview image; an operation (250) of identifying whether at least a portion of the first text has disappeared within the preview image according to a preset condition; and an operation (260) of removing the first visual object when at least a portion of the first text has disappeared according to the preset condition. and may include an operation (270) of maintaining the first visual object on the preview image when at least a part of the first text has not disappeared according to the preset condition.

[0184] According to one embodiment, the method may include an operation (260) of removing the first visual object when at least a part of the first text disappears according to the preset condition.

[0185] According to one embodiment, the operation of identifying whether at least a portion of the first text has disappeared according to the preset condition may include the operation (530) of identifying whether at least a portion of the first text has disappeared within the preview image due to the motion of the electronic device.

[0186] According to one embodiment, the operation of identifying whether at least a portion of the first text has disappeared according to the preset condition may include the operation (540) of identifying whether the first text has disappeared within the preview image as the first object containing the first text moves.

[0187] According to one embodiment, the operation of generating a first visual object including the first translation text comprises: applying the preview image, the first text, and the first translation text to an artificial intelligence model trained to generate the first visual object; and obtaining the first visual object from the artificial intelligence model; wherein the background color of the first visual object and the color of the first translation text may be determined by the artificial intelligence model.

[0188] According to one embodiment, as the first text is obscured by the movement of another object within the preview image, the method may include an operation of changing at least one of the background color of the first visual object or the color of the first translation text based on the color of the other object.

[0189] According to one embodiment, the method may include: an operation of predicting whether the first text will be obscured by another object as another object moves within the preview image; and an operation of setting the first visual object to be displayed on the top layer as it is predicted that the first text will be obscured.

[0190] According to one embodiment, the first visual object obtained from the artificial intelligence model includes additional information related to the first text, and the additional information may be information retrieved based on keywords within the first text.

[0191] According to one embodiment, the method may include: generating a second visual object comprising the second translation text based on the preview image, the second text within the preview image, and the second translation text for the second text; and overlaying the second visual object onto the second text within the preview image.

[0192] According to one embodiment, the method may include an operation of determining the overlapping sequence of the first visual object and the second visual object based on depth information of the first object including the first text and depth information of the second object including the second text.

[0193] According to one embodiment, as at least one of the first object or the second object moves within the preview image, the method may include: identifying that the first object and the second object overlap at least partially; and displaying the first visual object and the second visual object overlapping according to the overlap order as the first object and the second object overlap at least partially.

[0194] FIG. 14 is a block diagram of an electronic device (1401) in a network environment (1400) according to various embodiments. Referring to FIG. 14, in the network environment (1400), the electronic device (1401) may communicate with an electronic device (1402) through a first network (1498) (e.g., a short-range wireless communication network) or with an electronic device (1404) or a server (1408) through a second network (1499) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1401) may communicate with the electronic device (1404) through a server (1408). According to one embodiment, the electronic device (1401) may include a processor (1420), memory (1430), input module (1450), sound output module (1455), display module (1460), audio module (1470), sensor module (1476), interface (1477), connection terminal (1478), haptic module (1479), camera module (1480), power management module (1488), battery (1489), communication module (1490), subscriber identification module (1496), or antenna module (1497). In some embodiments, at least one of these components (e.g., connection terminal (1478)) may be omitted from the electronic device (1401), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (1476), camera module (1480), or antenna module (1497)) may be integrated into a single component (e.g., display module (1460)).

[0195] The processor (1420) can, for example, execute software (e.g., program (1440)) to control at least one other component (e.g., hardware or software component) of the electronic device (1401) connected to the processor (1420) and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (1420) can store commands or data received from other components (e.g., sensor module (1476) or communication module (1490)) in volatile memory (1432), process the commands or data stored in volatile memory (1432), and store the resulting data in non-volatile memory (1434). According to one embodiment, the processor (1420) may include a main processor (1421) (e.g., a central processing unit or an application processor) or an auxiliary processor (1423) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (1401) includes a main processor (1421) and an auxiliary processor (1423), the auxiliary processor (1423) may be configured to use less power than the main processor (1421) or to be specialized for a specified function. The auxiliary processor (1423) may be implemented separately from the main processor (1421) or as part thereof.

[0196] The auxiliary processor (1423) may control at least some of the functions or states associated with at least one component of the electronic device (1401) (e.g., display module (1460), sensor module (1476), or communication module (1490)) on behalf of the main processor (1421) while the main processor (1421) is in an inactive (e.g., sleep) state, or together with the main processor (1421) while the main processor (1421) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (1423) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (1480) or communication module (1490)). According to one embodiment, the auxiliary processor (1423) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (1401) itself where the artificial intelligence is performed, or through a separate server (e.g., server (1408)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0197] The memory (1430) can store various data used by at least one component of the electronic device (1401) (e.g., processor (1420) or sensor module (1476)). The data may include, for example, input data or output data for software (e.g., program (1440)) and related commands. The memory (1430) may include volatile memory (1432) or non-volatile memory (1434).

[0198] The program (1440) may be stored as software in memory (1430) and may include, for example, an operating system (1442), middleware (1444), or an application (1446).

[0199] The input module (1450) can receive commands or data to be used for a component of the electronic device (1401) (e.g., processor (1420)) from outside the electronic device (1401) (e.g., user). The input module (1450) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0200] The sound output module (1455) can output a sound signal to the outside of the electronic device (1401). The sound output module (1455) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0201] The display module (1460) can visually provide information to an external (e.g., user) of the electronic device (1401). The display module (1460) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (1460) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0202] The audio module (1470) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (1470) can acquire sound through the input module (1450) or output sound through the sound output module (1455) or an external electronic device (e.g., electronic device (1402)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (1401).

[0203] The sensor module (1476) can detect the operating state of the electronic device (1401) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (1476) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0204] The interface (1477) may support one or more specified protocols that can be used for the electronic device (1401) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (1402)). According to one embodiment, the interface (1477) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0205] The connection terminal (1478) may include a connector through which the electronic device (1401) can be physically connected to an external electronic device (e.g., electronic device (1402)). According to one embodiment, the connection terminal (1478) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0206] The haptic module (1479) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (1479) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0207] The camera module (1480) can capture still images and video. According to one embodiment, the camera module (1480) may include one or more lenses, image sensors, image signal processors, or flashes.

[0208] The power management module (1488) can manage the power supplied to the electronic device (1401). According to one embodiment, the power management module (1488) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0209] The battery (1489) can supply power to at least one component of the electronic device (1401). According to one embodiment, the battery (1489) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0210] The communication module (1490) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (1401) and an external electronic device (e.g., electronic device (1402), electronic device (1404), or server (1408)), and the performance of communication through the established communication channel. The communication module (1490) may include one or more communication processors that operate independently of the processor (1420) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1490) may include a wireless communication module (1492) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (1494) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (1404) via a first network (1498) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (1499) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1492) can identify or authenticate the electronic device (1401) within a communication network such as the first network (1498) or the second network (1499) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (1496).

[0211] The wireless communication module (1492) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (1492) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (1492) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (1492) can support various requirements specified in the electronic device (1401), external electronic device (e.g., electronic device (1404)), or network system (e.g., second network (1499)). According to one embodiment, the wireless communication module (1492) can support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.

[0212] An antenna module (1497) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (1497) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (1497) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (1498) or a second network (1499), may be selected from the plurality of antennas, for example, by a communication module (1490). A signal or power may be transmitted or received between the communication module (1490) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (1497).

[0213] According to various embodiments, the antenna module (1497) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0214] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0215] According to one embodiment, commands or data may be transmitted or received between the electronic device (1401) and an external electronic device (1404) through a server (1408) connected to a second network (1499). Each of the external electronic devices (1402, or 1404) may be the same or a different type of device as the electronic device (1401). According to one embodiment, all or part of the operations performed on the electronic device (1401) may be performed on one or more of the external electronic devices (1402, 1404, or 1408). For example, if the electronic device (1401) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (1401) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (1401). The electronic device (1401) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (1401) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1404) may include an Internet of Things (IoT) device. The server (1408) may be an intelligent server using machine learning and / or neural networks.According to one embodiment, an external electronic device (1404) or server (1408) may be included within the second network (1499). The electronic device (1401) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0216] According to one embodiment, the electronic device (1401) of FIG. 14 may correspond to the electronic device (1000) of FIG. 1 to FIG. 13, and the electronic device (1401) may perform the operation of the electronic device (1000) of FIG. 1 to FIG. 13.

[0217] A processor (1420) according to one embodiment can acquire a preview image through a camera (1480). The processor (1420) can activate the camera (1480) of the electronic device (1401) by running a camera application, and can acquire a preview image through the activated camera (1480). The processor (1420) can acquire a preview image in real time through the camera (1480) and display it on the screen of the electronic device (1401).

[0218] A processor (1420) according to one embodiment may obtain a first translated text by translating a first text within a preview image. The first text may be text that is the subject of translation, for example, text displayed on a specific object within the preview image. Additionally, the first translated text may be a translation result obtained by translating the first text. For example, the processor (1420) may identify the first text from the preview image using an optical character recognition (OCR) function and input the identified first text into an artificial intelligence model trained for translation, but is not limited thereto. In this case, the artificial intelligence model for translation may be an artificial intelligence model that performs translation based on text. Alternatively, for example, the processor (1420) may input the preview image into an artificial intelligence model trained for translation and obtain a translation result output from the artificial intelligence model. In this case, the artificial intelligence model for translation may be an artificial intelligence model that recognizes and translates text within an image.

[0219] A processor (1420) according to one embodiment may generate a first visual object including a first translation text based on a preview image, a first text, and a first translation text. According to one embodiment, the first visual object may be an object generated to include the first translation text and displayed on the screen of an electronic device (1401). For example, the first visual object may be generated to include the first translation text and a background. For example, the first visual object may be generated to include an image within the preview image, the first translation text, and a background. For example, the first visual object may be generated in the form of a window that is displayed separately from the preview image. For example, the first visual object may include a first translation text having a determined size and color, and may have a determined shape, size, and background color. For example, the background color of the first visual object may have the same or similar color as the color of the first object containing the first text. For example, the color of the first translation text within the first visual object may have a color distinct from the background color.

[0220] According to one embodiment, the processor (1420) may apply a preview image together with at least one of a first text or a first translation text to a first artificial intelligence model trained for generating a visual object. The first artificial intelligence model may be, for example, an artificial intelligence model trained to generate first visual objects having colors that are easy to visually recognize by a user, taking into account the color of the first text included in the preview image, the color of the first object on which the first text is displayed, and the color of other objects within the preview image. For example, the first artificial intelligence model may be a multimodal model that learns and processes relationships between data of various types or various modalities, such as text and images. The first artificial intelligence model may be a large multimodal model (LMM) trained using text data and image data. The first artificial intelligence model may generate a visual object containing a translation text based on a text prompt, an image prompt, or a prompt consisting of text and an image input to the first artificial intelligence model.

[0221] According to one embodiment, the processor (1420) may acquire a first visual object generated by taking into account the user's personal information. For example, the processor (1420) may acquire a first visual object by taking into account information regarding the user's eyesight and / or information regarding the user's color blindness. In this case, the processor (1420) may acquire a first visual object generated by taking into account the user's visual characteristics by applying information regarding the user's visual characteristics and a preview image together with at least one of a first text or a first translation text to a first artificial intelligence model. In this case, the first artificial intelligence model may be an artificial intelligence model trained to generate first visual objects having colors that are easy to visually recognize by the user by taking into account, for example, the color of the first text included in the preview image, the color of the first object on which the first text is displayed, the color of other objects within the preview image, and the user's visual characteristics.

[0222] A processor (1420) according to one embodiment may display a first visual object by overlapping it on a preview image. The processor (1420) may display the first visual object on the first text so as to overlap at least partially with the first text. The processor (1420) may display the first visual object so as to overlap at least partially with a first object containing the first text.

[0223] According to one embodiment, the processor (1420) can display the first visual object so that the first visual object is superimposed on the first text in real time by monitoring the position of the first text. In this case, the processor (1420) can identify the position of the first text within the preview image in real time and can move the display position of the first visual object in real time as the first text moves within the preview image.

[0224] According to one embodiment, the processor (1420) can render the preview image and the first visual object so that the preview image and the first visual object are displayed superimposed.

[0225] A processor (1420) according to one embodiment can identify whether at least a portion of the first text has disappeared from the preview image according to a preset condition. According to one embodiment, the processor (1420) can identify whether at least a portion of the first text has disappeared from the preview image according to the user's intention.

[0226] According to one embodiment, the processor (1420) can identify whether at least a portion of the first text has disappeared from the preview image according to the active intent of the user. For example, the processor (1420) can identify whether at least a portion of the first text has disappeared from the preview image due to the movement of the electronic device (1401). For example, the user can move the electronic device (1401) to change the shooting direction of the electronic device (1401), and as the direction in which the camera (1480) of the electronic device (1401) faces changes, at least a portion of the first text may disappear within the preview image. In this case, the processor (1420) can identify whether at least a portion of the first text has disappeared from the preview image due to the movement of the electronic device (1401) based, for example, on at least one of the distance moved by the electronic device (1401), the direction of movement, and the degree to which the shooting direction of the electronic device (1401) has changed.

[0227] According to one embodiment, the processor (1420) can identify whether at least a portion of the first text has disappeared from the preview image due to the user's passive intention. For example, the processor (1420) can identify whether at least a portion of the first text has disappeared from the preview image due to the movement of the first object. For example, when the first object, which is a subject in the real world, moves, the user may not change the shooting direction of the electronic device (1401). Accordingly, the position of the first object within the preview image may move, and the first object may disappear within the preview image. In this case, the processor (1420) can identify whether at least a portion of the first text has disappeared from the preview image due to the movement of the first object by, for example, monitoring the relative position of the first object and another object within the preview image.

[0228] According to one embodiment, the processor (1420) compares the degree of movement of the electronic device (1401) with a first threshold and the degree of movement of the first object with a second threshold, thereby allowing the processor (1420) to identify whether at least a portion of the first text has disappeared from the preview image according to a preset condition. For example, if the degree of movement of the electronic device (1401) (e.g., distance moved, direction of movement, amount of rotation, direction of rotation, etc.) is greater than the first threshold, or if the degree of movement of the first object (e.g., distance moved, direction of movement within the preview image, etc.) is greater than the second threshold, the processor (1420) can identify whether at least a portion of the first text has disappeared from the preview image according to a preset condition, but is not limited thereto.

[0229] According to one embodiment, if at least a portion of the first text disappears from the preview image according to a preset condition, the processor (1420) in operation 260 may remove the first visual object on the preview image. If at least a portion of the first text disappears from the preview image according to a preset condition, the processor (1420) may determine that the first text has disappeared from the preview image according to the user's intention and may not display the first translation text on the preview image. For example, the first text may disappear from the preview image by the user changing the direction of the camera (1480) of the electronic device (1401) to a different direction, in which case the processor (1420) may stop rendering the first visual object containing the first translation text for the first text. For example, the first text may disappear from the preview image by the first object containing the first text moving on its own, in which case the processor (1420) may stop rendering the first visual object containing the first translation text for the first text.

[0230] According to one embodiment, in the case where at least a portion of the first text has not disappeared from the preview image according to a preset condition, the processor (1420) in operation 270 may maintain the first visual object on the preview image. For example, the first object containing the first text may be obscured by another object, and in this case, it may be determined that the first text is not being displayed in the preview image regardless of the user's intention. In this case, even while the first text is not being displayed in the preview image, the processor (1420) may maintain the display of the first visual object containing the first translated text.

[0231] FIG. 15 is a drawing showing a system including a generative artificial intelligence model according to one embodiment.

[0232] Referring to FIG. 15, the User Query / Response Interface (2310) can receive user input. The user input may be in the form of natural language, images, and / or videos. Additionally, context information may be transmitted along with the user input. Context information may include various additional information at the time of user input. For example, information about the application currently being used by the user or the user's location information. Furthermore, user input may be in a mixed form of the aforementioned natural language, images, sounds, and context information. Additionally, user input may be in a non-natural language form, such as selecting a menu. The User Query / Response Interface (2310) can output results from a generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of actions requested by the user.

[0233] The AI ​​framework (2320) can receive input from the user and coordinate and control each component necessary to perform the user's intent based on the user's query.

[0234] User input received from the User Query / Response Interface (2310) can be sent to the Prompt design component (2321). The Prompt design component (2321) can be used to generate a prompt suitable for inputting user input into a Large Language Model (LLM) or Large Multimodal Models (LMM). The Prompt design component (2321) may be an AI component that uses machine learning algorithms or neural networks to develop better prompts over time. Based on user input, the Prompt design component (2321) can generate a prompt by accessing a knowledge component (e.g., knowledge repositories (2340)) containing user preference data, a prompt library, and prompt examples, and can pass the generated prompt to the LLM or LMM.

[0235] The API / Plug-in management component (2323) can perform the role of communicating with external information when there is a request for additional information when passing user input as input to a generative model. The API / Plug-in management component (2323) establishes a channel to communicate with the outside of the AI ​​Interface via the API, and through the established channel, it can enable access to various data sources (e.g., knowledge repositories (2340)). Additionally, if the application or service needs to perform an action that executes the user input as a final step rather than an intermediate result, the API / Plug-in management component (2323) can request that action from the application / service component (2330) via the API. The information obtained from the outside can be used to generate a prompt in the Prompt design component (2321) along with the user input, or it can be passed as input to the generative model.

[0236] The Refiner component (e.g., output modification component (2325)) allows for detailed tuning of the output from a generative model. For instance, the Refiner component can verify whether the content generated by LLM and / or LMM is irrelevant, contains biased content, or includes harmful content. Additionally, the Refiner component can determine the extent to which the output matches the user's desired outcome and, if necessary, proceed with additional processing. Furthermore, the Refiner component can configure and provide hints to the user to help avoid unwanted outputs.

[0237] A Generative AI Model (2350) can generally refer to an artificial intelligence neural network that generates new forms of data based on user input information. A Generative AI Model (2350) may include a model that generates images and / or a model that generates language. Models that generate images include, but are not limited to, GANs (generative adversarial networks) and VAEs (variational autoencoders), and examples include Diffusion-based generative models that use VAEs and Transformer structures. Models that generate language are models trained to output the most statistically appropriate output value based on input values, and examples include models such as CHAT-GPT 3 and CHAT-GPT 4. There are also LMMs that can recognize various forms of data input, such as text, images, and voice, and generate new data corresponding to them.

[0238] According to one embodiment, the electronic device (1000) of FIGS. 1 to 13 and the electronic device (1401) of FIG. 14 may be configured to include at least some of the User Query / Response Interface (2310) AI framework (2320), application / service component (2330), knowledge repositories (2340), or Generative AI Model (2350) of FIG. 15. According to one embodiment, at least some of the User Query / Response Interface (2310) AI framework (2320), application / service component (2330), knowledge repositories (2340), or Generative AI Model (2350) of FIG. 15 may be included in another electronic device (e.g., another user's electronic device and / or server).

[0239] According to one embodiment, an electronic device (1000, 1401) that provides a translation result for text within a preview image comprises: a camera (1480); a display (1460); and a memory (1430) that stores commands; The electronic device may provide a translation result for text in a preview image, wherein, when the instructions are executed individually or collectively by the at least one processor, the electronic device may: acquire a preview image (210) through the camera, acquire a first translated text (220) by translating a first text in the preview image, create a first visual object (230) including the first translated text based on the preview image, the first text, and the first translated text, and display the preview image and the first visual object (240), wherein the first visual object is overlaid on the first text in the preview image, identify (250) whether at least part of the first text has disappeared in the preview image according to a preset condition, and if at least part of the first text has not disappeared according to the preset condition, maintain (270) the first visual object on the preview image.

[0240] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be able to: remove the display of the first visual object (260) when at least a part of the first text disappears according to the preset condition.

[0241] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may identify whether at least a portion of the first text has disappeared within the preview image due to motion of the electronic device.

[0242] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may: identify whether the first text disappears within the preview image as the first object containing the first text moves.

[0243] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device is configured to: apply the preview image, the first text, and the first translation text to an artificial intelligence model trained to generate the first visual object, and obtain the first visual object from the artificial intelligence model, and the background color of the first visual object and the color of the first translation text may be determined by the artificial intelligence model.

[0244] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be able to: change at least one of the background color of the first visual object or the color of the first translation text based on the color of the other object as the first text is obscured by the movement of the other object within the preview image.

[0245] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may: predict whether the first text will be obscured by the other object as the other object moves within the preview image, and, as it is predicted that the first text will be obscured, set the first visual object to be displayed on the top layer.

[0246] According to one embodiment, the first visual object obtained from the artificial intelligence model includes additional information related to the first text, and the additional information may be information retrieved based on keywords within the first text.

[0247] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may: generate a second visual object including the second translation text based on the preview image, the second text within the preview image, and the second translation text for the second text, and overlay the second visual object on the second text within the preview image.

[0248] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may determine the overlapping sequence of the first visual object and the second visual object based on depth information of the first object containing the first text and depth information of the second object containing the second text.

[0249] According to one embodiment, a computer-readable recording medium may be provided that records a program for executing a method comprising: acquiring a preview image obtained through a camera; acquiring a first translated text by translating a first text within the preview image; generating a first visual object including the first translated text based on the preview image, the first text, and the first translated text; displaying the preview image and the first visual object, wherein the first visual object is overlaid on the first text within the preview image; identifying whether at least a portion of the first text has disappeared within the preview image according to a preset condition; removing the display of the first visual object when at least a portion of the first text has disappeared according to the preset condition; and maintaining the first visual object on the preview image when at least a portion of the first text has not disappeared according to the preset condition.

[0250] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0251] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0252] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0253] Various embodiments of the present document may be implemented as software (e.g., program (1440)) comprising one or more instructions stored in a storage medium (e.g., internal memory (1436) or external memory (1438)) readable by a machine (e.g., electronic device (1401)). For example, a processor (e.g., processor (1420)) of the machine (e.g., electronic device (1401)) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0254] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0255] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. A method in which an electronic device provides a translation result for text within a preview image, The operation of acquiring a preview image through a camera; The operation of obtaining a first translated text by translating the first text within the above preview image; The operation of generating a first visual object including the first translation text based on the above preview image, the above first text, and the above first translation text; An operation in which the above preview image and the above first visual object are displayed, wherein the above first visual object is overlaid on the above first text within the above preview image; An operation to identify whether at least a portion of the first text has disappeared within the preview image according to a preset condition; and The operation of maintaining the first visual object on the preview image when at least a portion of the first text has not disappeared according to the preset condition; A method including 2. In Paragraph 1, An operation to remove the first visual object when at least a portion of the first text disappears according to the preset condition; A method that further includes.

3. In Paragraph 1, The operation of identifying whether at least a portion of the first text has disappeared according to the preset condition is, A method comprising identifying whether at least a portion of the first text has disappeared within the preview image due to motion of the electronic device.

4. In Paragraph 1, The operation of identifying whether at least a portion of the first text has disappeared according to the preset condition is, A method comprising identifying whether the first text disappears within the preview image as the first object containing the first text moves.

5. In Paragraph 1, The operation of creating a first visual object including the first translation text above is, The operation of applying the above preview image, the above first text, and the above first translation text to an artificial intelligence model trained to generate the above first visual object; and The operation of obtaining the first visual object from the above artificial intelligence model; Includes, A method in which the background color of the first visual object and the color of the first translation text are determined by the artificial intelligence model.

6. In Paragraph 5, An action of changing at least one of the background color of the first visual object or the color of the first translation text based on the color of the other object as the first text is obscured by the movement of the other object within the preview image; A method including 7. In Paragraph 5, An operation to predict whether the first text will be obscured by the other object as the other object moves within the preview image; and An action of setting the first visual object to be displayed on the top layer as it is predicted that the first text will be obscured; A method that includes 8. In Paragraph 5, The first visual object obtained from the artificial intelligence model includes additional information related to the first text or at least one image within the preview image, and The above additional information is information retrieved based on keywords within the first text.

9. In Paragraph 1, The operation of generating a second visual object including the second translation text based on the above preview image, the second text within the above preview image, and the second translation text for the second text; and The operation of overlaying the second visual object onto the second text within the preview image; A method including 10. In Paragraph 9, An operation to determine the overlapping sequence of the first visual object and the second visual object based on depth information of the first object including the first text and depth information of the second object including the second text; A method including 11. In Paragraph 10, As at least one of the first object or the second object moves within the above preview image, an action of identifying that the first object and the second object overlap at least partially; and An operation of displaying the first visual object and the second visual object by overlapping them according to the overlapping order, as the first object and the second object overlap at least partially; A method including 12. An electronic device that provides a translation result for text within a preview image, camera; display; Memory for storing instructions; and At least one processor; Includes, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: A preview image is obtained through the above camera, and By translating the first text within the above preview image, a first translated text is obtained, and Based on the above preview image, the above first text, and the above first translation text, a first visual object including the above first translation text is created, and Display the above preview image and the above first visual object, wherein the above first visual object is overlaid on the above first text within the above preview image, and Identify whether at least a portion of the first text above has disappeared within the preview image according to preset conditions, and An electronic device that maintains the first visual object on the preview image when at least a portion of the first text is not lost according to the preset condition.

13. In Paragraph 12, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: An electronic device that removes the display of the first visual object when at least a portion of the first text disappears according to the preset condition.

14. In Paragraph 12, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: An electronic device that identifies whether at least a portion of the first text above has disappeared within the preview image due to motion of the electronic device.

15. The operation of acquiring a preview image obtained through a camera; The operation of obtaining a first translated text by translating the first text within the above preview image; The operation of generating a first visual object including the first translation text based on the above preview image, the above first text, and the above first translation text; An operation in which the above preview image and the above first visual object are displayed, wherein the above first visual object is overlaid on the above first text within the above preview image; An operation to identify whether at least a portion of the first text has disappeared within the preview image according to a preset condition; An operation to remove the display of the first visual object when at least a portion of the first text disappears according to the preset condition; and The operation of maintaining the first visual object on the preview image when at least a portion of the first text has not disappeared according to the preset condition; A computer-readable recording medium that records a program for executing a method including