Electronic device and method for providing translation function, and storage medium

An AI-driven OCR process translates text within images while preserving layout, addressing the challenge of language translation in electronic devices, ensuring accurate and contextually preserved results.

WO2026010199A1PCT designated stage Publication Date: 2026-01-08SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/008202
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-27
Filing Date
2025-06-13
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing electronic devices struggle to efficiently translate text and images within web pages across different languages, particularly in a manner that preserves the original layout and context.

Method used

The implementation of an AI-driven optical character recognition (OCR) process to identify and translate text objects within images, followed by inpainting the translated text objects back into the original image layout, ensuring context preservation.

Benefits of technology

Enables accurate and contextually preserved translation of text and images within web pages, enhancing user experience by maintaining the original layout and readability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025008202_08012026_PF_FP_ABST
    Figure KR2025008202_08012026_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device may comprise: a memory storing instructions; one or more processors; and a display. The instructions, when executed by the at least one processor, may cause the electronic device to: display a first representation including a first image of a web page; identify an input for translating a first language into a second language; perform optical character recognition (OCR) of a first image for obtaining first text objects of the first language from the first image; perform the translation on the first text objects to obtain second text objects of the second language; identify arrangement information of the second text objects on the basis of arrangement information of the first text objects; inpaint the second text objects onto the first image by using the arrangement information of the second text objects, so as to generate a second image including the second text objects; and display a second representation of the web page including the second image.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and storage medium providing translation function

[0001] The descriptions below relate to electronic devices, methods, and storage media that provide translation functions.

[0002] Electronic devices can provide various types of content. For example, the electronic device may execute an application that provides a web browser. For example, based on the execution of the application, the electronic device may display a web page of the web browser on the display of the electronic device. For example, the web page may include text and / or images.

[0003] The above information may be provided as background information to aid in understanding the present disclosure. None of the above is claimed to be prior art related to the present disclosure or can be used to determine prior art related to the present disclosure.

[0004] In embodiments of the present disclosure, an electronic device is provided. The electronic device may include a memory comprising one or more storage media and storing instructions. The electronic device may include at least one processor comprising a processing circuit. The electronic device may include a display. The instructions, when individually and / or collectively executed by the at least one processor, may cause the electronic device to display a first representation of a web page via the display. The first representation may include a first image. The instructions, when individually and / or collectively executed by the at least one processor, may cause the electronic device to identify an input for performing a translation from a first language to a second language with respect to the web page. The instructions, when individually and / or collectively executed by the at least one processor, may cause the electronic device to perform optical character recognition (OCR) of the first image to obtain first text objects in the first language from the first image based on the input. Placement information of the first text objects within the first image may be further identified by performing the OCR of the first image. The instructions, when individually and / or collectively executed by the at least one processor, may cause the electronic device to perform the translation from the first language to the second language on the obtained first text objects, thereby obtaining second text objects in the second language.The instructions, when individually and / or collectively executed by the at least one processor, may cause the electronic device to identify placement information of the second text objects in a second image to be generated based on the placement information of the first text objects. The instructions, when individually and / or collectively executed by the at least one processor, may cause the electronic device to generate a second image including the second text objects by inpainting the second text objects into the first image using the placement information of the second text objects. The instructions, when individually and / or collectively executed by the at least one processor, may cause the electronic device to display, via the display, a second representation of the web page including the second image.

[0005] In embodiments of the present disclosure, a method performed by an electronic device is provided. The method performed by the electronic device may include an operation of displaying a first representation of a web page. The first representation may include a first image. The method may include an operation of identifying an input for performing a translation from a first language to a second language with respect to the web page. The method may include an operation of performing optical character recognition (OCR) on the first image to obtain first text objects in the first language from the first image based on the input. Placement information of the first text objects within the first image may be further identified by performing the OCR on the first image. The method may include an operation of performing the translation from the first language to the second language with respect to the obtained first text objects, thereby obtaining second text objects in the second language. The method may include an operation of identifying, based on the placement information of the first text objects, placement information of the second text objects within a second image to be generated. The method may include an action of generating the second image including the second text objects by inpainting the second text objects onto the first image using the arrangement information of the second text objects. The method may include an action of displaying a second representation of the web page including the second image.

[0006] In embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs comprising instructions that, when individually or collectively executed by at least one processor of an electronic device including a display, cause the electronic device to display a first representation of a web page through the display. The first representation may include a first image. The non-transitory computer-readable storage medium may store one or more programs comprising instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to identify an input for performing a translation from a first language to a second language with respect to the web page. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to perform optical character recognition (OCR) of the first image to obtain first text objects in the first language from the first image based on the input. Information about the arrangement of the first text objects within the first image may be further identified by performing the OCR of the first image. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to perform the translation from the first language to the second language on the obtained first text objects, thereby obtaining second text objects in the second language.The non-transitory computer-readable storage medium may store one or more programs comprising instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to identify arrangement information of the second text objects in a second image to be generated based on the arrangement information of the first text objects. The non-transitory computer-readable storage medium may store one or more programs comprising instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to generate a second image including the second text objects by inpainting the second text objects into the first image using the arrangement information of the second text objects. The non-transitory computer-readable storage medium may store one or more programs comprising instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to display, through the display, a second representation of the web page including the second image.

[0007] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.

[0008] Figure 2 illustrates an example of a model driven by an electronic device.

[0009] Figure 3a illustrates an example of a neural network structure.

[0010] Figure 3b illustrates an exemplary block diagram of an electronic device.

[0011] Figures 4a and 4b illustrate comparative examples of methods for translating text and / or images of a web page.

[0012] Figure 5 illustrates an example of a method for translating text contained in an image of a web page based on an AI (artificial intelligence) model.

[0013] Figure 6 illustrates an example of a flow of operations for translating elements of a web page based on an AI model.

[0014] Figure 7a illustrates an example of a method for requesting translation of texts and images among elements of a web page based on an AI model.

[0015] FIG. 7b illustrates an example of a method for requesting translation of images in a first area of ​​a web page and images in a second area of ​​a web page based on an AI model.

[0016] FIG. 8a illustrates an example of a method for determining whether an image on a web page contains a text object based on an AI model.

[0017] Fig. 8b illustrates an example of a method for obtaining text information of a text object included in an image when the image of a web page includes a text object, based on an AI model.

[0018] Figure 8c illustrates an example of a method for generating a translated image using image and text information based on an AI model.

[0019] Figure 9 illustrates an example of a method for displaying a web page using a resource map that stores images and translated images.

[0020] Figure 10 illustrates an example of a method for generating a translated image using another web page related to a web page based on an AI model and displaying a web page including the translated image.

[0021] Figures 11a and 11b illustrate examples of web pages provided upon activation of the translation function.

[0022] Figure 12 illustrates an example of how to perform translation for a portion of a web page.

[0023] Figure 13 illustrates an example of how to convert a translated element of a web page into an element before translation.

[0024] Figure 14 illustrates an example of a flowchart for a method of generating a translated image using image and text information based on an AI model, upon determining that an image on a web page contains text.

[0025] The terms used in this disclosure are used only to describe specific embodiments and may not be intended to limit the scope of other embodiments. The singular expression may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by those of ordinary skill in the art described in this disclosure. Terms defined in general dictionaries among the terms used in this disclosure may be interpreted as having the same or similar meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined in this disclosure. In some cases, even if a term is defined in this disclosure, it cannot be interpreted to exclude embodiments of the present disclosure.

[0026] The various embodiments of the present disclosure described below illustrate a hardware-based approach as an example. However, since the various embodiments of the present disclosure include techniques utilizing both hardware and software, the various embodiments of the present disclosure do not exclude a software-based approach.

[0027] In addition, in the present disclosure, expressions such as "more than" or "less than" may be used to determine whether a specific condition is satisfied or fulfilled, but this is merely a description for expressing an example and does not exclude descriptions such as "more than" or "less than." A condition described as "more than" may be replaced with "more than," a condition described as "less than" may be replaced with "less than," and a condition described as "more than and less than" may be replaced with "more than and less than." In addition, hereinafter, "A" to "B" mean at least one of the elements from A (including A) to B (including B). hereinafter, "C" and / or "D" mean at least one of "C" or "D," that is, including {"C", "D", "C" and "D"}. hereinafter, the meaning of "about E" may be replaced with a value within a margin of error of ±5% or ±10% based on E.

[0028] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.

[0029] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0030] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0031] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0032] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0033] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0034] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0035] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0036] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0037] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0038] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0039] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0040] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0041] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0042] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0043] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).

[0044] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0045] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0046] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0047] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0048] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0049] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0050] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0051] Figure 2 illustrates an example of a model driven by an electronic device.

[0052] The electronic device (101) of FIG. 2 may be an example of the electronic device (101) of FIG. 1. Referring to FIG. 2, according to one embodiment, the electronic device (101) may include at least one of a central processing unit (CPU) (210), a neural processing unit (NPU) (220), a graphic processing unit (GPU) (230), or a memory (130). The CPU (210), the NPU (220), the GPU (230), and the memory (130) may be electrically and / or operably coupled with each other by an electronic component, such as a communication bus (205). The type and / or number of hardware components included in the electronic device (101) are not limited to those illustrated in FIG. 2. Hereinafter, the hardware components being operatively coupled may mean that a direct connection or an indirect connection between the hardware components is established, either wired or wireless, such that a second hardware component is controlled by a first hardware component among the hardware components.

[0053] Referring to FIG. 2, according to one embodiment, an electronic device (101) may include hardware components (e.g., a CPU (210), an NPU (220), a GPU (230), and / or a memory (130)) for performing operations on a model (240) related to an artificial neural network. The model (240) and / or the artificial neural network may include a recognition model implemented in software or hardware that mimics the computational capabilities of a biological system by using a large number of artificial neurons (or nodes). For example, the electronic device (101) according to one embodiment may, based on the model (240), perform functions similar to human cognitive operations or learning processes. Based on calculations indicated by the model (240) and performed in a chain by a plurality of parameters, the electronic device (101) may output data including generalized information about input data. The memory (130) of the electronic device (101) may store the plurality of parameters related to the model (240). The CPU (210), NPU (220), and / or GPU (230) of the electronic device (101) may include circuitry for performing the calculations that are performed serially by the plurality of parameters.

[0054] According to one embodiment, the CPU (210) of the electronic device (101) may include a hardware component for processing data based on one or more instructions. The hardware component for processing data may include, for example, an arithmetic and logic unit (ALU), a floating point unit (FPU), and / or a field programmable gate array (FPGA). In one embodiment, the CPU (210) may be referred to as an application processor (AP). The number of CPUs (210) may be one or more. For example, the CPU (210) may have a multi-core processor structure such as a dual core, a quad core, or a hexa core. The CPU (210) of FIG. 2 may be an example of the processor (120) and / or the main processor (121) of FIG. 1.

[0055] According to one embodiment, the NPU (220) of the electronic device (101) may include hardware components dedicated to computations related to the model (240). For example, the NPU (220) may include a plurality of circuits for performing computations (e.g., multiplication and / or addition) sequentially and / or in parallel based on the model (240). The plurality of circuits included in the NPU (220) may be referred to as neural engines. The NPU (220) may perform the computations based on a designated data type (e.g., floating point number and / or integer) related to the model (240).

[0056] According to one embodiment, the GPU (230) of the electronic device (101) may include one or more pipelines that perform multiple operations for executing instructions related to computer graphics and / or parallel computing. For example, the pipeline of the GPU (230) may include a graphics pipeline or a rendering pipeline for generating a three-dimensional image and generating a two-dimensional raster image from the generated three-dimensional image. By using the graphics pipelines, calculations related to artificial neural networks can be executed substantially simultaneously.

[0057] The CPU (210), NPU (220), and GPU (230) of FIG. 2 may be included as different integrated circuits in the electronic device (101), or may be included in a single integrated circuit (single IC) based on a system on chip (SoC). For example, the CPU (210), the NPU (220), the GPU (230), or a combination thereof may be included in a single integrated circuit included in the electronic device (101). The type of processing unit included based on the SoC is not limited to the above example, and for example, other hardware components (e.g., a communication processor) not shown in FIG. 2 may be included in a single integrated circuit together with the CPU (210), the NPU (220), and the GPU (230). Hereinafter, in terms of the subject of the calculations of the artificial neural network directed by the model (240), the CPU (210), the NPU (220), the GPU (230), or a combination thereof may be referred to as an AI (artificial intelligence) accelerator (or accelerator). The AI ​​accelerator may be referred to as an accelerator.

[0058] A memory (130) of an electronic device (101) according to one embodiment may include a hardware component for storing data and / or instructions input and / or output to a CPU (210), an NPU (220), and / or a GPU (230). The memory (130) may include, for example, a volatile memory (132) such as a random-access memory (RAM) and / or a non-volatile memory (134) such as a read-only memory (ROM). The volatile memory (132) may include, for example, at least one of a dynamic RAM (DRAM), a static RAM (SRAM), a cache RAM, and a pseudo SRAM (PSRAM). The non-volatile memory (134) may include, for example, at least one of a programmable ROM (PROM), an erasable ROM (EPROM), an electrically erasable ROM (EEPROM), a flash memory, a hard disk, a compact disk, and an embedded multi media card (eMMC). The memory (130), the volatile memory (132), and the non-volatile memory (134) of FIG. 2 may correspond to the memory (130), the volatile memory (132), and the non-volatile memory (134) of FIG. 1, respectively.

[0059] Within the memory (130), one or more instructions (or commands) that indicate operations to be performed by the CPU (210), the NPU (220), and / or the GPU (230) based on data may be stored. A set of one or more instructions may be referred to as firmware, an operating system, a process, a routine, a sub-routine, and / or an application. For example, the CPU (210), the NPU (220), and / or the GPU (230) of the electronic device (101) may perform at least one of the operations of FIGS. 1 and 2 when a set of a plurality of instructions distributed in the form of an operating system, firmware, a driver, and / or an application is executed. Hereinafter, the fact that an application is installed in an electronic device (101) may mean that one or more instructions provided in the form of an application are stored in the memory (130) of the electronic device (101), and that the one or more applications are stored in a format (e.g., a file having an extension specified by an operating system of the electronic device (101)) executable by the CPU (210), NPU (220), and / or GPU (230) of the electronic device (101). The one or more pipelines of FIG. 2 may be configured to be executed by the CPU (210) of the electronic device (101) and to control the CPU (210), NPU (220), and / or GPU (230).

[0060] In one embodiment, the electronic device (101) can identify a model (240) based on one or more files stored in the non-volatile memory (134). The one or more files can be related to an application (e.g., an application (146) of FIG. 1), a middleware (e.g., a middleware (144) of FIG. 1), and / or an operating system (e.g., an operating system (142) of FIG. 1) installed in the electronic device (101). In one embodiment where a model (240) related to an application is stored in the non-volatile memory (134), the CPU (210) can identify the model (240) stored in the non-volatile memory (134) based on execution of the application. Identifying a model (240) stored in the non-volatile memory (134) may include copying (or loading) a plurality of parameters associated with the model (240), stored in the non-volatile memory (134), into the volatile memory (132). Identifying a model (240) stored in the non-volatile memory (134) may include obtaining a plurality of instructions for performing calculations indicated by the model (240), based on the plurality of parameters stored in the volatile memory (132). Accelerators, such as a CPU (210), an NPU (220), and / or a GPU (230), may execute one or more functions associated with the model (240) indicated by the plurality of parameters stored in the volatile memory (132), based on the plurality of instructions. The above one or more functions may include at least one of a function for performing training of a model (240), a function for performing inference on input data based on the model (240), a function for performing image-based object recognition, voice recognition, and / or handwriting recognition using the trained model (240), and a function personalized to a user of the electronic device (101) based on a neural network. However, the embodiment is not limited thereto.

[0061] According to one embodiment, the electronic device (101) may perform calculations directed by the model (240) based on a plurality of parameters associated with the model (240). The plurality of parameters may include weights assigned to a plurality of nodes and / or connections between the plurality of nodes indicated by the model (240). The plurality of parameters may include hyperparameters related to the model (240). The hyperparameters may include, for example, at least one of a learning rate, a cost function, a regularization parameter, a mini-batch size, a number of training iterations, a number of hidden layers, a metaparameter, or a free parameter.

[0062] According to one embodiment, the electronic device (101) may perform calculations related to input data based on the model (240) using an accelerator. The electronic device (101) may obtain output data from the input data input to the model (240) based on performing chained (or serial or consecutive) calculations based on a plurality of parameters of the model (240). The input data may include a plurality of numeric values ​​that are preprocessed to be input to the model (240). The plurality of numeric values ​​may indicate a vector to be input to the model (240). The electronic device (101) may obtain at least one numeric value indicating output data by modifying the plurality of numeric values ​​included in the input data based on a plurality of parameters and calculations indicated by the model (240). The calculations related to the model (240) and / or the plurality of parameters may be distinguished by an operation, a graph, and / or a layer.

[0063] Referring to FIG. 2, an exemplary sequence of operations (245-1, 245-2, 245-3, 245-4) included in a model (240) is illustrated. The operations (245-1, 245-2, 245-3, 245-4) may be referred to as sub-models included in the model (240). The CPU (210) of the electronic device (101) may select the model (240) and / or the sub-model from among models corresponding to different sizes, depending on the size of the input data, and load the model (240) and / or the sub-model to the NPU (220) and / or the GPU (230).

[0064] Each of the operations (245-1, 245-2, 245-3, 245-4) may include a group of calculations that the electronic device (101) sequentially performs based on the operation of the model (240). The operations (245-1, 245-2, 245-3, 245-4) may be distinguished by the type of calculation performed by the electronic device (101). Referring to FIG. 2, the electronic device (101) may sequentially perform calculations on input data based on the order of the operations (245-1, 245-2, 245-3, 245-4) within the model (240). For example, operation (245-1) may indicate one or more convolution calculations based on a convolution filter associated with the model (240). Operation (245-2) performed after operation (245-1) may instruct one or more depthwise convolution calculations related to the model (240) on the input data changed by operation (245-1). Operation (245-3) performed after operation (245-2) may instruct mean pooling calculations related to the model (240) on the input data changed by operations (245-1, 245-2). Operation (245-4) performed after operation (245-3) may instruct convolution calculations related to the model (240) on the input data changed by operations (245-1, 245-2, 245-3, 245-4).

[0065] In one embodiment, numerical values ​​included in input data input to the model (240) may be changed based on connections between a plurality of nodes included in the model (240). The plurality of nodes may be divided into units of layers. In one embodiment, where the plurality of parameters include weights connecting two nodes of different layers of the model (240), the electronic device (101) may apply weights to values ​​corresponding to nodes of a specific layer to obtain values ​​corresponding to nodes of another layer connected to the specific layer. Within the model (240), a layer including nodes into which values ​​included in the input data are input may be referred to as an input layer. The last layer among the sequentially connected layers within the model (240) may be referred to as an output layer. Each of the operations (245-1, 245-2, 245-3, and 245-4) of FIG. 2 may include at least one of the layers included in the model (240). For example, operation (245-1) may indicate a group of interconnected layers based on a convolution filter among the layers within the model (240). Hereinafter, a graph may mean a graph formed by connections between nodes included in the layers. The graph included in the model (240) may be distinguished by operations (245-1, 245-2, 245-3, 245-4) included in the model (240).

[0066] Figure 3a illustrates an example of a neural network structure.

[0067] In the description with respect to FIG. 3a, the operation of the model or layer can be understood as the operation of a central processing unit (CPU) (e.g., CPU (210) of FIG. 2) and / or a neural processing unit (NPU) (e.g., NPU (220) of FIG. 2).

[0068] For example, the first layer (L1) may be a convolution layer, and the second layer (L2) may be a sampling layer. The artificial neural network model may further include an activation layer and further include layers that perform other types of operations.

[0069] Each of the plurality of layers can receive input image data or a feature map generated from a previous layer as an input feature map, and generate an output feature map by operating the input feature map. At this time, the feature map may mean data in which various characteristics of the input data are expressed. The feature maps (FM1, FM2, FM3) may have, for example, a two-dimensional matrix or a three-dimensional matrix form. The feature maps (FM1 to FM3) have a width (W) (or referred to as a column), a height (H) (or referred to as a row), and a depth (D), which may correspond to the x-axis, y-axis, and z-axis on the coordinate system, respectively. At this time, the depth (D) may be referred to as the number of channels.

[0070] The first layer (L1) can generate a second feature map (FM2) by convolving the first feature map (FM1) with a weight map (WM). The weight map (WM) can filter the first feature map (FM1) and may be referred to as a filter or a kernel. For example, the depth of the weight map (WM), i.e., the number of channels, is the same as the depth of the first feature map (FM1), e.g., the number of channels, and the same channels of the weight map (WM) and the first feature map (FM1) can be convolved. The weight map (WM) can be shifted in a manner that traverses the first feature map (FM1) using a sliding window. The amount of shifting can be referred to as a "stride length" or "stride." During each shift, each of the weights included in the weight map (WM) can be multiplied and added to all feature values ​​in the area overlapping the first feature map (FM1). As the first feature map (FM1) and the weight map (WM) are convolved, one channel of the second feature map (FM2) can be generated. Although one weight map (WM) is shown in Fig. 3a, in reality, multiple weight maps can be convolved with the first feature map (FM1) to generate multiple channels of the second feature map (FM2). In other words, the number of channels of the second feature map (FM2) can correspond to the number of weight maps.

[0071] The second layer (L2) can generate the third feature map (FM3) by changing the spatial size of the second feature map (FM2). For example, the second layer (L2) can be a sampling layer. The second layer (L2) can perform up-sampling or down-sampling, and the second layer (L2) can select some of the data included in the second feature map (FM2). For example, a two-dimensional window (WD) can be shifted on the second feature map (FM2) in units of the size of the window (WD) (e.g., a 4 * 4 matrix), and a value at a specific location (e.g., 1 row, 1 column) in an area overlapping the window (WD) can be selected. The second layer (L2) can output the selected data as data of the third feature map (FM3). As another example, the second layer (L2) can be a pooling layer. In this case, the second layer (L2) may select the maximum value (or average value of feature values) of the feature values ​​in the area overlapping the window (WD) in the second feature map (FM2). The second layer (L2) may output the selected data as data of the third feature map (FM3).

[0072] Accordingly, a third feature map (FM3) with a changed spatial size can be generated from the second feature map (FM2). The number of channels of the third feature map (FM3) and the number of channels of the second feature map (FM2) may be the same. Meanwhile, according to an exemplary embodiment of the present disclosure, the operation speed of the sampling layer may be faster than that of the pooling layer, and the sampling layer may improve the quality of the output image (e.g., in terms of Peak Signal to Noise Ratio (PSNR)). For example, the operation by the pooling layer may take longer to operate than the operation by the sampling layer because the maximum value or the average value must be calculated.

[0073] Depending on the embodiment, the second layer (L2) may not be limited to a sampling layer or a pooling layer. For example, the second layer (L2) may be a convolution layer similar to the first layer (L1). The second layer (L2) may convolve the second feature map (FM2) with a weight map to generate a third feature map (FM3). In this case, the weight map on which the convolution operation is performed in the second layer (L2) may be different from the weight map (WM) on which the convolution operation is performed in the first layer (L1).

[0074] An Nth feature map can be generated from an Nth layer through a plurality of layers including a first layer (L1) and a second layer (L2). The Nth feature map can be input to a reconstruction layer located at the back end of an artificial neural network model from which output data is output. The reconstruction layer can generate an output image based on the Nth feature map. In addition, the reconstruction layer can receive a plurality of feature maps, such as a first feature map (FM1) and a second feature map (FM2), in addition to the Nth feature map, and generate an output image based on the plurality of feature maps.

[0075] For example, the restoration layer may be a convolution layer or a deconvolution layer. Depending on the embodiment, it may also be implemented as another type of layer capable of restoring an image from a feature map.

[0076] Figure 3b illustrates an exemplary block diagram of an electronic device.

[0077] Referring to FIG. 3B, according to one embodiment, an electronic device (101) may include a processor (310), a display (320), and a memory (330). However, the embodiments of the present disclosure are not limited thereto. For example, the processor (310), the display (320), and the memory (330) may be electrically and / or operably coupled with each other by a communication bus. Hereinafter, operably coupled hardware components may mean that a direct connection or an indirect connection is established between the hardware components, either wired or wireless, such that a second hardware component is controlled by a first hardware component among the hardware components. Although illustrated based on different blocks, the embodiment is not limited thereto, and some of the hardware components illustrated in FIG. 3b (e.g., at least a portion of the processor (310) and the memory (330)) may be included in a single integrated circuit such as a system on a chip (SoC) or a system in package (SIP). The type and / or number of hardware components included in the electronic device (101) is not limited to those illustrated in FIG. 3b. For example, the electronic device (101) may include only some of the hardware components illustrated in FIG. 3b.

[0078] According to one embodiment, the processor (310) of the electronic device (101) may include a hardware component for processing data based on one or more instructions. The hardware component for processing data may include, for example, an arithmetic and logic unit (ALU), a floating point unit (FPU), and a field programmable gate array (FPGA). As an example, the hardware component for processing data may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing (DSP), a microcontroller (MCU), and / or a neural processing unit (NPU). The number of processors (310) may be one or more. For example, the processor (310) may have a multi-core processor structure such as a dual core, a quad core, or a hexa core. The processor (310) of FIG. 3B may be substantially identically applied to the processor (120) of FIG. 1.

[0079] For example, the processor (310) may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including in the claims, may include various processing circuits including at least one processor, one or more of which may be configured to individually and / or collectively perform the various functions described below in a distributed manner. As used herein, when "processor," "at least one processor," and "one or more processors" are described as being configured to perform various functions, these terms encompass, for example, and without limitation, situations where one processor performs some of the recited functions and other processor(s) perform other parts of the recited functions, and also situations where one processor may perform all of the recited functions. Additionally, the at least one processor may include a combination of processors that perform the various functions enumerated / disclosed, for example, in a distributed manner. At least one processor may execute program instructions to achieve or perform the various functions.

[0080] According to one embodiment, a display (320) of an electronic device (101) can output visualized information to a user of the electronic device (101). For example, the display (320) can be controlled by a processor (310) including circuits such as a CPU, a GPU (graphic processing unit), and / or a DPU (display processing unit) to output visualized information to the user. The display (320) can include a flexible display, a flat panel display (FPD), and / or electronic paper. The display (320) can include a liquid crystal display (LCD), a plasma display panel (PDP), and / or one or more light emitting diodes (LEDs). The LEDs can include organic LEDs (OLEDs). Embodiments are not limited thereto, and for example, if the electronic device (101) includes a lens for transmitting external light (or ambient light), the display (320) may include a projector (or projection assembly) for projecting light onto the lens. In one embodiment, the display (320) may be referred to as a display panel and / or a display module. For example, the display (320) of FIG. 3B may be an example of the display module (160) of FIG. 1. A display (320) that supports a touch function may be referred to as a touch screen. The display (320) may further include a structure capable of detecting an input using a stylus pen, such as an electro-magnetic resonance (EMR) or an active electrostatic solution (AES).For example, at least some of the events described in this document (e.g., events or inputs for translating a web page) may be triggered using a stylus pen. For example, a user may designate a portion of content displayed on the display (320) that he or she wishes to translate through a stylus pen-based input (e.g., drawing around at least a portion of the desired area). For example, if a user draws a circular shape on content (e.g., an image, text, video, document, or handwriting input) using a stylus pen and handwrites the word “translate” around it, the electronic device (101) may identify this as a request to translate the selected content. For example, the user may manually input various commands (e.g., specifying a language to translate into, requesting a search) in addition to “translate.”

[0081] According to one embodiment, the memory (330) of the electronic device (101) may include a hardware component for storing data and / or instructions input to and / or output from the processor (310). The memory may include, for example, a volatile memory such as a random-access memory (RAM), and / or a non-volatile memory such as a read-only memory (ROM). The volatile memory may include, for example, at least one of a dynamic RAM (DRAM), a static RAM (SRAM), a cache RAM, and a pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a flash memory, a hard disk, a compact disc, and an embedded multimedia card (eMMC). The specific details of the memory (330) of FIG. 3b can be applied substantially identically to the details of the memory (130) of FIG. 1.

[0082] According to one embodiment, one or more instructions (or commands) representing operations and / or actions to be performed on data by the processor (310) of the electronic device (101) may be stored in the memory (330) of the electronic device (101). A set of one or more instructions may be referred to as a program, firmware, an operating system, a process, a routine, a sub-routine, and / or an application. Hereinafter, when an application is installed in the electronic device (e.g., the electronic device (101)), it may mean that one or more instructions provided in the form of an application are stored in the memory (330), and that the one or more applications are stored in a format executable by the processor of the electronic device (e.g., a file having an extension designated by the operating system of the electronic device (101)). According to one embodiment, the electronic device (101) may execute one or more instructions stored in the memory (330) to perform the operations of FIGS. 6 and 14. For example, the one or more instructions, when executed by the processor (310), may cause the electronic device (101) to perform at least some of the operations of FIGS. 6 and 14.

[0083] According to one embodiment, the memory (330) may include (or store) an AI model (340). The AI ​​model (340) of FIG. 3b may be an example of the model (240) of FIG. 2.

[0084] For example, the memory (330) may temporarily or non-temporarily store data related to the AI ​​model (340). For example, the memory (330) may store data related to the AI ​​model (340), such as layers, weights, and operations of the AI ​​model (340), and may update the data related to the AI ​​model (340) based on the output value of the AI ​​model (340). For example, the memory (330) may store the output value of the AI ​​model (340).

[0085] An AI model (340) according to one embodiment may be an artificial neural network model written in a specified language and including a plurality of layers and / or operations. The AI ​​model (340) according to one embodiment may be at least one of various types of networks, such as a convolution neural network (CNN), a region with convolution neural network (R-CNN), a region proposal network (RPN), a recurrent neural network (RNN), a stacking-based deep neural network (S-DNN), a state-space dynamic neural network (S-SDNN), a deconvolution network, a deep belief network (DBN), a restricted Boltzmann machine (RBM), a fully convolutional network, a long short-term memory (LSTM) network, and a classification network. The AI ​​model (340) according to one embodiment may be trained on specified data, acquire input data, and perform operations based on the input data to generate output data.

[0086] An AI model (340) according to one embodiment may include an input layer, a hidden layer, and an output layer. The input layer may be related to input values ​​input to the AI ​​model (340). The hidden layer may perform a MAC operation (multiply-accumulate) and an activation operation on the input values ​​to output a feature map. The output layer may be related to the result value of the operation performed in the hidden layer.

[0087] In one embodiment, the AI ​​model (340) may be stored in the memory (330). In one embodiment, operations based on the AI ​​model (340) may be performed in a processor (310) (e.g., a central processing unit (CPU) and / or a neural processing unit (NPU)).

[0088] Hereinafter, the present disclosure exemplifies operations (or functions) performed using an AI model (340) stored in an electronic device (101) and executed on the electronic device (101), but the present disclosure is not limited thereto. For example, at least some of the operations (or functions) of the present disclosure may be performed based on an external electronic device (or an AI model within the external electronic device) connected to the electronic device (101).

[0089] For example, the processor (310) can perform operations based on an AI model (340) stored in the memory (330).

[0090] A processor (310) according to one embodiment may include a compiler (315).

[0091] A compiler (315) according to one embodiment may compile source code written in a specific language into object code that can be processed by a target program and / or target hardware. For example, the compiler (315) may load an AI model (340) from memory (330) and compile the AI ​​model (340) into a binary usable by the processor (310).

[0092] According to one embodiment, a processor (310) can check an activation function included in a compiled AI model (340).

[0093] For example, the processor (310) may activate the first function and / or the second function in response to including the first type function in the activation function of the compiled AI model (340).

[0094] In one embodiment, the first type function may be a Rectified Linear Unit (ReLU) function. The ReLU function is a function that outputs 0 if the input value is negative, and outputs the input value as is if the input value is positive.

[0095] According to one embodiment, the processor (310) may perform a MAC (multiply-accumulate) operation and an activation operation on input values ​​in a hidden layer to output a feature map. For example, the MAC operation may be an operation that multiplies input values ​​by corresponding weights and then sums the multiplied values. The activation operation may be an operation that inputs the result of the MAC operation into an activation function and outputs a result value.

[0096] According to one embodiment, the processor (310) may skip operations on designated values ​​included in the input of the hidden layer. For example, the first function may be a zero-skipping function. The zero-skipping function may be a function that skips MAC operations on 0s included in input values ​​(e.g., feature maps, result values ​​of previous hidden layers). For example, the second function may be a feature map compression function. The feature map compression function may be a function in which the processor extracts non-zero values ​​from the output values ​​of the hidden layer (e.g., feature maps) and compresses the feature map.

[0097] An AI model (340) (not shown) according to one embodiment may include an input layer, a hidden layer, and an output layer. For example, the input layer is a layer related to input values ​​input to the AI ​​model (340). For example, the hidden layer may perform a MAC (multiply-accumulate) operation and an activation operation on input values ​​to output a feature map. For example, the MAC operation may be an operation that multiplies input values ​​by corresponding weights and adds the multiplied values. For example, the activation operation may be an operation that inputs the result of the MAC operation to an activation function and outputs a result value. The activation function may be of various types. For example, the activation function may include, but is not limited to, a sigmoid function, a tangent function, a Relu function, a Riki-Relu function, a Max-Out function, and / or an Elu function. For example, the hidden layer may be composed of at least one layer. For example, if the hidden layer is composed of a first hidden layer and a second hidden layer, the first hidden layer may perform a MAC operation and an activation operation based on the input value of the input layer to output a feature map, and the feature map, which is the result value of the first hidden layer, may become the input value of the second hidden layer. The second hidden layer may perform a MAC operation and an activation operation based on the feature map, which is the result value of the first hidden layer. For example, the output layer may be a layer related to the result value of the operation performed in the hidden layer.

[0098] In the description related to FIG. 3b, at least some of the parts depicted as being included in the processor (310) may be SW modules executed in the processor (310), and in this case, the operations performed by the parts may be understood as operations of the processor (310).

[0099] Although not illustrated in FIG. 3B, the electronic device (101) may include a communication circuit (not illustrated). For example, the electronic device (101) may communicate with an external electronic device or a server using the communication circuit. For example, the communication circuit may include hardware for supporting transmission and / or reception of electrical signals. The communication circuit may include, for example, at least one of a modem, an antenna, and an optical / electronic (O / E) converter. The communication circuit may support transmission and / or reception of electrical signals based on various types of communication means, such as Ethernet, Bluetooth, Bluetooth low energy (BLE), ZigBee, long term evolution (LTE), and 5G new radio (NR). Specific details regarding the communication circuit of FIG. 3B may be substantially identical to the communication module (190) and / or antenna module (197) of FIG. 1.

[0100] In addition, although not illustrated in FIG. 3B, the electronic device (101) may further include an output device. For example, the output device may include an output means for outputting information in a form other than the visualized information provided through the display (320). For example, the output device may include a speaker for outputting an acoustic signal. For example, the speaker may be used to provide auditory information. For example, the output device may include a motor for providing haptic feedback based on vibration. For example, the motor may be used to provide tactile information. For example, the output device may include the acoustic output module (155) and / or the haptic module (179) of FIG. 1.

[0101] For example, the AI ​​model (340) according to one embodiment may be referred to as a trained model, an on-device AI model, or an AI model within the device, as an AI model stored within the electronic device (101). For example, the AI ​​model (340) may include a plurality of AI engines. For example, the plurality of AI engines may include a first AI engine and a second AI engine. For example, the first AI engine may be referred to as a text extraction engine or a text extraction library. For example, the second AI engine may be referred to as a translation engine or a translation image generation engine.

[0102] For example, the first AI engine may be used to determine whether an image contains text. For example, the first AI engine may analyze an input image to determine whether the image contains text. As a result of the analysis of the image, the first AI engine may output a value indicating that the image contains text or a value indicating that the image does not contain text. For specific details related thereto, reference may be made to FIG. 8A below.

[0103] For example, the first AI engine may obtain text information of the text within the image upon determining that the image contains text. For example, the text information may include an optical character recognition (OCR) result of the text within the image (or contents or characters of the text), location information (or coordinates) of the text within the image, or size information of a virtual block (or text box) in which the text is located. However, the present disclosure is not limited thereto. For example, the text information may further include rotation information indicating a degree to which the text within the image is rotated with respect to a specified reference. For specific details related thereto, reference may be made to FIG. 8B below.

[0104] For example, the second AI engine may be used to perform a translation on text. For example, the second AI engine may output text translated into a target language using text in an input source language. For example, the text in the source language input to the second AI engine may include text having text attributes within the configuration information of a web page, or text included in an image based on the first AI engine. In one example, the second AI engine may generate a modified image including the translated text using text information of an image and the text within the image. For specific details related thereto, reference may be made to FIGS. 8B to 8C below.

[0105] As a non-limiting example, the target language may include the target language of the most recently performed translation. As a non-limiting example, the target language may be determined based on a priority order of languages ​​preset for the electronic device (101). As a non-limiting example, the target language may be determined based on a usage pattern of translations performed by the electronic device (101) (or a user's translation usage pattern). For example, the usage pattern may include translations performed for each web page. In other words, the usage pattern may vary for each web page.

[0106] In the example of FIG. 3B, the electronic device (101) is illustrated as including a single AI model (340), but the present disclosure is not limited thereto. For example, the AI ​​model (340) may be implemented as multiple AI models. For example, the first AI engine may be implemented with the functionality of the first AI model, and the second AI engine may be implemented with the functionality of the second AI model.

[0107] For example, the electronic device (101) can execute an application stored in the memory (330). For example, the application may include a software application (or a web browser application) that provides a web browser. As the electronic device (101) executes the application, the electronic device (101) can display a web page of the web browser on the display (320). For example, the electronic device (101) can display a user interface (UI) including at least a portion of the web page on the display (320). For example, the UI may be referred to as a screen, a display area, a display portion, a representation, or a screen interface.

[0108] For example, the electronic device (101) can execute a translation function of the web page. For example, the electronic device (101) can use configuration information defining elements (or web contents) of the web page to identify elements having text properties and elements having image properties among the elements. For example, the element having the text property can be identified based on a tag (or text tag) indicating text of the configuration information. For example, the element having the image property can be identified based on a tag (or image tag) indicating an image of the configuration information. For example, the configuration information can include code information used to configure the web page. For example, the configuration information can include HTML (hypertext markup language). The configuration information can be referenced as a markup language file of the web page. Examples of the configuration information can be referenced in the table below.

[0109] [Correction pursuant to Rule 91, July 16, 2025]

[0110] For example, the electronic device (101) can perform a translation on an element having a text attribute among the elements identified using the configuration information. Alternatively, for example, the electronic device (101) can transmit an element having an image attribute among the elements of the web page identified using the configuration information to an external electronic device (or server), and receive a translated text of the text in the element, thereby displaying the translated text above the image. In other words, the translated text can be displayed by at least partially overlapping (or floating, overlapping, or overlaying) the element. For specific details related thereto, reference may be made to FIGS. 4A and 4B below.

[0111] Figures 4a and 4b illustrate comparative examples of methods for translating text and / or images of a web page.

[0112] FIGS. 4A and 4B illustrate comparative examples (400, 405, 407) of a method in which an electronic device (101) displays a UI of a web page on a display (320) and provides a translation function of the web page.

[0113] FIG. 4a illustrates a comparative example (400) in which a UI (410) is displayed in a state in which the translation function of a web page is deactivated, and a comparative example (405) in which a UI (410) is displayed in a state in which the translation function of a web page is activated.

[0114] Referring to comparative example (400), the electronic device (101) can display a UI (410) including at least a portion of the web page on the display (320). The UI including at least a portion of the web page of the present disclosure can be referred to as a representation of the web page. For example, the UI (410) can include a plurality of icons (421, 422, 423) and an image (430). For example, the UI (410) can include a plurality of texts (421a, 422a, 423a) corresponding to the plurality of icons (421, 422, 423).

[0115] For example, the icon (421) may have a shape representing (or associating) news. The text (421a) may include text in a first language (e.g., Korean) representing 'news'. For example, the icon (422) may have a shape representing (or associating) shopping. The text (422a) may include text in a first language (e.g., Korean) representing 'shopping'. For example, the icon (423) may have a shape representing (or associating) a map. The text (423a) may include text in a first language (e.g., Korean) representing 'map'. For example, the image (430) may include a plurality of texts (431, 432, 433). For example, the text (431) in the image (430) may include text representing '10 billion live lactic acid bacteria'. For example, text (432) within image (430) may include text indicating 'bio core'. For example, text (433) within image (430) may include text indicating 'see more'.

[0116] For example, the configuration information for the web page related to the UI (410) of FIG. 4a may include information about icons (421, 422, 423) and images (430), which are elements having image properties. For example, the configuration information for the web page may include information about texts (421a, 422a, 423a), which are elements having text properties. For example, the configuration information may not include information about texts (431, 432, 433) included in the images (430).

[0117] For example, the electronic device (101) may identify an event (or input) for translating the web page. For example, the event may include an input to an icon representing a translation function included in a menu of the UI (410). However, the present disclosure is not limited thereto. For example, the event may include a designated gesture. Or, for example, the event may include an input to a pop-up message suggesting activation of the translation function.

[0118] For example, the electronic device (101) may perform a translation of the web page upon identifying the event and display a UI (410) including at least a portion of the translated web page. Referring to comparative example (405), the electronic device (101) may display a UI (410) including at least a portion of the translated web page on the display (320).

[0119] For example, the UI (410) of the comparative example (405) may include a plurality of icons (421, 422, 423) and an image (430). For example, the UI (410) may include a plurality of texts (421b, 422b, 423b) corresponding to the plurality of icons (421, 422, 423). Compared to the comparative example (400), the texts (421b, 422b, 423b) of the comparative example (405) may be texts translated into a target language. For example, the text (421b) may include text (or character) in a second language (e.g., English) indicating 'News'. For example, the text (422b) may include text in a second language (e.g., English) indicating 'Shopping'. For example, text (423b) may include text in a second language (e.g., English) representing 'Map'. In the example of Fig. 4a, the first language may be referred to as a source language, and the second language may be referred to as a target language. In the example of Fig. 4a, for convenience of explanation, it is assumed that the source language is different from the target language, but the present disclosure is not limited thereto. For example, if the text in the UI (410) is text in the second language, the text in the second language may not be translated even if the translation function is activated.

[0120] For example, when the translation function is activated, the electronic device (101) may perform translation for elements having text properties within the web page, and refrain from (or may not perform, stop, skip) performing translation for elements having image properties within the web page. Referring to the comparative example (405), the image (430) may include a plurality of text objects (431, 432, 433). For example, the text object (431) within the image (430) may include text indicating '10 billion healthy probiotics'. For example, the text object (432) within the image (430) may include text indicating 'bio core'. For example, the text object (433) within the image (430) may include text indicating 'view more'. In other words, the image (430) of the comparative example (400) and the image (430) of the comparative example (405) may be substantially identical.

[0121] FIG. 4B illustrates a comparative example (400) in which a UI (410) is displayed in a state in which the translation function of a web page is deactivated, and a comparative example (407) in which a UI (410) is displayed in a state in which the translation function of the web page is activated. The comparative example (400) of FIG. 4B may be substantially identical to the comparative example (400) of FIG. 4A. Accordingly, the specific details of the comparative example (400) of FIG. 4B may be substantially identical to the details of the comparative example (400) of FIG. 4A.

[0122] For example, the electronic device (101) may perform translation of the web page upon identifying an event for translation of the web page and display a UI (410) including at least a portion of the translated web page. For example, the electronic device (101) may perform translation of elements having text properties and elements having image properties among elements of the web page upon identifying the event. For example, the electronic device (101) may provide information (or an image) about an element having an image property to an external electronic device (or a server) connected to the electronic device (101) in order to perform translation of an element having an image property. For example, the electronic device (101) may receive information about translated text from text in an element having an image property from the external electronic device (or a server). For example, the electronic device (101) may display the translated text on an element having an image property. Referring to comparative example (407), the electronic device (101) can display a UI (410) including at least a portion of the translated web page on the display (320).

[0123] For example, the UI (410) of the comparative example (407) may include a plurality of icons (421, 422, 423) and an image (440). For example, the UI (410) may include a plurality of texts (421b, 422b, 423b) corresponding to the plurality of icons (421, 422, 423). Compared to the comparative example (400), the texts (421b, 422b, 423b) of the comparative example (407) may be texts translated into a target language. For example, the text (421b) may include text (or character) in a second language (e.g., English) indicating 'News'. For example, the text (422b) may include text in a second language (e.g., English) indicating 'Shopping'. For example, text (423b) may include text in a second language (e.g., English) representing 'Map'. In the example of Fig. 4b, the first language may be referred to as a source language, and the second language may be referred to as a target language. In the example of Fig. 4b, for convenience of explanation, it is assumed that the source language is different from the target language, but the present disclosure is not limited thereto. For example, if the text in the UI (410) is text in the second language, the text in the second language may not be translated even if the translation function is activated.

[0124] For example, the electronic device (101) may perform translation for elements having image properties within the web page when the translation function is activated. Referring to comparative example (407), the text box (440) (or text block) may include a plurality of texts (441, 442, 443). For example, the text (441) within the text box (440) may include text indicating '10 billion live lactobacillus'. For example, the text (442) within the text box (440) may include text indicating 'Bio core'. For example, the text (443) within the text box (440) may include text indicating 'Detailed view'. In the UI (410) of comparative example (407), the text box (440) may be displayed overlapping on the image (430). For example, the UI (410) of the comparative example (407) may include icons (421, 422, 423) and an image (430), and a text box (440) at least partially overlapped on the image (430) may be displayed. For example, the text box (440) may have a box (or text box) shape that displays text information received from the external electronic device.

[0125] As described above, when the translation function of a web page is activated, the electronic device (101) may perform translation only for elements having text properties among the elements of the web page, as in the comparative example (405) of FIG. 4a, or may translate text within an element having an image property among the elements and then display the translated text by overlapping it on the element having the image property, as in the comparative example (407) of FIG. 4b. In the case of the comparative example (405), since translation is not performed for elements having image properties, translation of the entire web page may not be provided, contrary to the user's intention. In the case of the comparative example (407), translation is performed for elements having image properties, but since an image including the translated text is displayed on the original image (or element), contrary to the user's intention, a damaged web page may be displayed, which may lower the user's understanding of the web page.

[0126] Hereinafter, the present disclosure can determine an image containing text among elements having image properties of a web page based on an AI model stored in an electronic device (101), and perform a translation on the image containing text. For example, the present disclosure can perform a translation on an image containing text based on the AI ​​model, and provide a naturally translated web page by adjusting the properties of the translated text to be similar to the properties of the image. In addition, the present disclosure can define an order (or method) of translation on elements of the web page performed based on the AI ​​model. Accordingly, the present disclosure can efficiently perform the translation performed based on the AI ​​model. In addition, the present disclosure can store a translated image generated by performing a translation based on the AI ​​model and a translated text included in the translated image in a separate storage space (or resource map) related to the web page. For example, the present disclosure can provide users with faster translated web pages by utilizing translated images and text stored in the separate storage space when a web page is changed or when the web page's translation function is activated or deactivated. For specific details on a method for performing web page translation based on an AI model, refer to FIG. 5 below.

[0127] Figure 5 illustrates an example of a method for translating text contained in an image of a web page based on an AI (artificial intelligence) model.

[0128] FIG. 5 illustrates a comparative example (400) in which a UI (410) is displayed in a state in which a translation function of a web page is deactivated, and an example (500) in which a UI (410) representing a web page translated based on an AI model (340) is displayed in a state in which the translation function of the web page is activated. The comparative example (400) of FIG. 5 may be substantially identical to the comparative example (400) of FIG. 4a. Accordingly, the specific details of the comparative example (400) of FIG. 5 may be substantially identically applied to the details of the comparative example (400) of FIG. 4a.

[0129] For example, the electronic device (101) may perform translation of the web page upon identifying an event for translation of the web page and display a UI (410) including at least a portion of the translated web page. For example, the electronic device (101) may perform translation of elements having text properties and elements having image properties among elements of the web page upon identifying the event. For example, the electronic device (101) may provide information (or images) about elements having image properties to an AI model stored (or included) in the electronic device (101) in order to perform translation of elements having image properties. For example, the electronic device (101) may generate a translated image (530) based on the stored (or included) AI model. For example, the translated image (530) may represent an image including text translated from text in an element having an image property. For example, the electronic device (101) can display a UI (410) including a translated image (530) on the display (320).

[0130] For example, the UI (410) of the example (500) may include a plurality of icons (421, 422, 423) and an image (530). For example, the UI (410) of the example (500) may include a plurality of texts (421b, 422b, 423b) corresponding to the plurality of icons (421, 422, 423). Compared to the comparative example (400), the texts (421b, 422b, 423b) of the example (500) may be texts translated into a target language. For example, the texts (421b, 422b, 423b) may be translated into the target language based on the AI ​​model (340). For example, the electronic device (101) can use the texts (421a, 422a, 423a) of the comparative example (400) as input and generate texts (421b, 422b, 423b) translated into the target language based on the AI ​​model (340). For example, the electronic device (101) can perform translation based on the AI ​​model (340) for elements having text properties prioritizing translation for elements having image properties. Specific details related to the order of requesting translation are exemplified and described below with reference to FIGS. 7A and 7B.

[0131] For example, text (421b) may include text (or characters) in a second language (e.g., English) indicating 'News'. For example, text (422b) may include text in a second language (e.g., English) indicating 'Shopping'. For example, text (423b) may include text in a second language (e.g., English) indicating 'Map'. In the example of FIG. 5, the first language may be referred to as a source language, and the second language may be referred to as a target language. In the example of FIG. 5, for convenience of explanation, it is assumed that the source language is different from the target language, but the present disclosure is not limited thereto. For example, if the text in the UI (410) is text in the second language, the text in the second language may not be translated even if the translation function is activated.

[0132] For example, the electronic device (101) may perform translation of elements having image properties within the web page based on the AI ​​model (340) when the translation function is activated. Referring to example (500), the image (530) may include a plurality of text objects (531, 532, 533). For example, the text object (531) within the image (530) may include a character representing '10 billion live lactobacillus'. For example, the text object (532) within the image (530) may include a character representing 'Bio core'. For example, the text object (533) within the image (530) may include a character representing 'Detailed view'.

[0133] For example, the electronic device (101) may determine whether the image (430) includes text based on the AI ​​model (340). For example, the electronic device (101) may determine whether the image (430) includes text (e.g., texts (431, 432, 433)) based on the first AI engine of the AI ​​model (340). For example, the electronic device (101) may output, based on the AI ​​model (340), information indicating that the image (430) includes text or information indicating that the image (430) does not include text. Specific details related thereto are exemplified and described below with reference to FIG. 8A.

[0134] For example, the electronic device (101) may obtain (or extract) text information about a text object in the image (430) based on determining that the image (430) includes a text object based on the AI ​​model (340). For example, the electronic device (101) may determine whether the image (430) is an image including a text object. For example, the text information may include content (or OCR result) of the text object in the image (430), position information (or arrangement information) of the text within the image (430), and size information of a virtual block in which the text is located. For example, the electronic device (101) may obtain the text information using the input image (430) based on the first AI engine of the AI ​​model (340). Specific details related thereto are exemplified and described below with reference to FIG. 8B.

[0135] For example, the electronic device (101) may generate an image (530) including a translated text object using the text information about the image (430) and the text object within the image (430) based on the AI ​​model (340). For example, the electronic device (101) may perform a translation on the text object (e.g., text objects (431, 432, 433)) within the image (430) based on the second AI engine of the AI ​​model (340). For example, the electronic device (101) may generate text objects (531, 532, 533) translated into the target language from the text objects (431, 432, 433) within the image (430) based on the second AI engine of the AI ​​model (340). For example, the electronic device (101) can obtain (or identify) text information of translated text objects (531, 532, 533) within an image (530) to be generated based on the text information of the text objects (431, 432, 433) within the image (430). For example, the electronic device (101) can generate an image (530) including translated text objects (531, 532, 533) based on the second AI engine of the AI ​​model (340). For example, the electronic device (101) can generate the image (530) by removing (or outpainting) text objects (431, 432, 433) from the image (430) and inputting (or inpainting) translated text objects (531, 532, 533) from the image (430).

[0136] For example, the position and properties (or styles) of text objects (531, 532, 533) in the image (530) can be adjusted based on the AI ​​model (340) (or the second AI engine of the AI ​​model (340)). For example, the electronic device (101) can adjust the properties (or styles) of the text objects (531, 532, 533) according to the representative colors of the pixels of the image (430) based on the AI ​​model (340). For example, the properties (or styles) can include at least one of the color, thickness, size, font, or background color of the text. By displaying an image (530) including text objects (531, 532, 533) whose properties are adjusted (or applied) according to the representative colors of the pixels of the image (430) within the UI (410), the translated image (530) can be displayed more naturally. Accordingly, the user's usability and understanding of the translated content can be improved.

[0137] In the example (500) of FIG. 5, a case is illustrated where a translated image (530) is generated based on an AI model (340) for one (a) image (430) included in the UI (410), but the present disclosure is not limited thereto. For example, the electronic device (101) can generate translated images including a text object translated based on the AI ​​model (340) for a plurality of images included in the UI (410), and display the translated images within the UI (410). Specific details regarding a method for generating a translated image are exemplified and described below with reference to FIG. 8c.

[0138] In the UI (410) of the example (500), the image (530) may be displayed by being changed from (or replaced with) the image (430). For example, the electronic device (101), when executing the application for the web browser, may store a resource map for one or more elements including the image (430) of the web page in a storage area used by the application. For example, the storage area may be included in at least a portion of the storage area of ​​the memory (330). For example, the electronic device (101), after generating the image (530) based on the AI ​​model, may store the image (530) together with the image (430) in the resource map. For example, when displaying the UI (410), the electronic device (101) may display the image (530) instead of the image (430) in the UI (410) by identifying the image (530) stored in the resource map. Specific details on how to display an image (530) identified within the above resource map within the UI (410) are exemplified and described below with reference to FIG. 9. The electronic device (101) can provide a translated web page to a user more quickly by utilizing the image (530) stored within the above resource map.

[0139] In FIG. 5, an example is described in which an electronic device (101) performs a translation for an element (e.g., an image (430)) having an image property among elements included in a web page based on an AI model (340) and generates a translated image, but the present disclosure is not limited thereto. For example, the processor (310) of the electronic device (101) may perform at least some of the operations (or functions) performed by the AI ​​model (340). As a non-limiting example, the processor (310) may identify an element having an image property using configuration information of a web page, and determine whether a text object is included in the element by performing OCR (optical character recognition) on the element. If a text object is included in the element, the processor (310) may extract the text object and perform a translation on the text object. When extracting the text object, the processor (310) can identify text information (e.g., layout information, font size, location) of the text object within the element, and can identify text information of the translated text object within the translated element (or image). The processor (310) can generate the translated element based on removing (or outpainting) the text object within the element before translation and adding (or inpainting) the translated text object within the element. Or, for example, the processor (310) can generate the translated image based on adding (or inpainting) translated texts above the text objects within the image before translation. At this time, the text objects before translation may not be visually visible due to the background color of the translated texts within the translated image.The background color may be adjusted according to the representative color of the image before translation. In other words, the processor (310) may generate the translated image by inpainting the translated texts for some of the areas where the text objects before translation are located in the image before translation, and inpainting the representative color of the first image for other areas. In the example, generating the translated element may be performed using the AI ​​model (340).

[0140] As illustrated in FIG. 5, specific details on a method for performing translation on elements having text properties (e.g., texts (421a, 422a, 423a)) and elements having image properties (e.g., images (430)) among elements included in a web page based on an AI model (340) are exemplified and described below with reference to FIG. 6.

[0141] Figure 6 illustrates an example of a flow of operations for translating elements of a web page based on an AI model.

[0142] At least some of the methods of FIG. 6 may be performed by the electronic device (101) of FIG. 3B. For example, at least some of the methods may be controlled by the processor (310) of the electronic device (101). In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0143] Referring to FIG. 6, in operation (600), the electronic device (101) may execute an application for a web browser. For example, the application may include a software application (or a web browser application) that provides a web browser. As the electronic device (101) executes the application, the electronic device (101) may display a web page of the web browser on the display (320). For example, the electronic device (101) may display a user interface (UI) including at least a portion of the web page on the display (320). For example, the UI may be referred to as a screen, a display area, a display portion, a representation, or a screen interface.

[0144] For example, the electronic device (101) may identify an event for translating the web page. For example, the event may include an input to an icon representing a translation function included in a menu of the UI. However, the present disclosure is not limited thereto. For example, the event may include a designated gesture. Or, for example, the event may include an input to a pop-up message suggesting activation of the translation function. For example, the electronic device (101) may execute the translation function of the web page based on identifying the event.

[0145] For example, the electronic device (101) can identify configuration information (or a markup language file) for the web page. For example, the electronic device (101) can obtain the configuration information based on executing the translation function of the web page. For example, the electronic device (101) can obtain the configuration information for the web page using a web application programming interface (API) (e.g., selection). In the above example, the case where the configuration information is obtained based on executing the translation function (or identifying the event) is described, but the present disclosure is not limited thereto. For example, the electronic device (101) can also obtain the configuration information based on executing the application. For example, the configuration information may include code information used to configure the web page. For example, the configuration information may include HTML (hypertext markup language).

[0146] In operation (605), the electronic device (101) may identify elements of the web page using configuration information about the web page. For example, the electronic device (101) may identify (or distinguish) elements having text properties (or text) and elements having image properties (or images) among the elements of the web page using the configuration information. In FIG. 6, for convenience of explanation, a case where multiple elements of the web page are identified using the configuration information is described, but the present disclosure is not limited thereto. For example, one (a) element of the web page may be identified using the configuration information. For example, the element having the text property may be identified based on a tag (or indicator) indicating text of the configuration information. For example, the element having the image property may be identified based on a tag indicating an image of the configuration information.

[0147] In operation (610), the electronic device (101) can determine whether the element is an image. For example, the electronic device (101) can perform operation (620) on an element having an image attribute among the elements identified using the configuration information. For example, the electronic device (101) can perform operation (615) on an element having a text attribute among the elements identified using the configuration information.

[0148] In operation (615), the electronic device (101) may perform a translation of text. For example, the electronic device (101) may perform a translation of an element having the text property among the elements of the web page based on the AI ​​model (340). For example, the electronic device (101) may perform a translation of an element having the text property input to the AI ​​model (340). For example, the language of the text of the element having the text property may be referred to as a source language. For example, the language of the text to be translated may be referred to as a target language.

[0149] For example, the electronic device (101) can identify the source language of the text of the element having the text attribute based on the AI ​​model (340). For example, the electronic device (101) can identify the language of the input text based on the AI ​​model (340). For example, if the source language of the text and the target language are different, the electronic device (101) can perform a translation of the text of the element based on the AI ​​model (340). Alternatively, if the source language of the text and the target language are different, the electronic device (101) can refrain from translating the text of the element.

[0150] In the above example, the electronic device (101) is illustrated as translating text based on the AI ​​model (340), but the present disclosure is not limited thereto. As a non-limiting example, the processor (310) of the electronic device (101) (or a program executed by the processor (310)) can perform the translation of text.

[0151] In operation (620), the electronic device (101) may determine whether the image is an image having a text object. For example, the electronic device (101) may determine, based on the AI ​​model (340), whether an element having the image property includes a text object. For example, the electronic device (101) may determine whether the image input to the AI ​​model (340) includes a text object. For example, the electronic device (101) may output a first value based on the AI ​​model (340) upon determining that the image includes a text object. For example, the electronic device (101) may output a second value based on the AI ​​model (340) upon determining that the image does not include a text object.

[0152] In the above example, the electronic device (101) is described as determining, based on the AI ​​model (340), whether the image is an image having the text object (or whether the image includes the text object), but the present disclosure is not limited thereto. As a non-limiting example, the processor (310) of the electronic device (101) (or a program executed by the processor (310)) can determine whether the image is an image having the text object (or whether the image includes the text object).

[0153] In operation (620), the electronic device (101) may perform operation (625) upon determining that the image does not have a text object. Alternatively, in operation (620), the electronic device (101) may perform operation (630) upon determining that the image has a text object.

[0154] In operation (625), the electronic device (101) may refrain from performing translation on the image. For example, the electronic device (101) may refrain from performing translation on the image based on the AI ​​model (340) upon determining that the image does not contain a text object (or upon determining that the image is not an image having a text object). Thereafter, when displaying the UI of the web page, the electronic device (101) may display the image within the UI without changing (or without translating) the image that does not contain a text object.

[0155] In operation (630), the electronic device (101) may obtain text information of the image. For example, the electronic device (101) may obtain text information of a text object in the image based on the AI ​​model (340) based on determining that the image includes a text object (or determining that the image is an image having a text object).

[0156] For example, the text information may include an OCR result of the text object within the image (or contents or characters of the text object), location information (or coordinates, arrangement information) of the text object within the image, or size information of a virtual block (or text box) in which the text object is positioned. However, the present disclosure is not limited thereto. For example, the text information may further include rotation information indicating a degree to which the text object within the image is rotated with respect to a specified reference.

[0157] In operation (635), the electronic device (101) may perform translation on an image and generate a translated image. For example, the electronic device (101) may perform translation on an image based on an AI model (340) using text information of an image and a text object within the image. In the example, the information provided to the AI ​​model (340) to perform translation on an image may include the image itself or a bitmap constituting the image. For example, performing translation on an image may include translating a text object within the image into text in a target language. For example, if the language (or source language) of a text object within the image is the same as the target language, translation on the corresponding text may not be performed. Alternatively, for example, if the language (or source language) of a text object within the image is different from the target language, translation on the corresponding text object may be performed.

[0158] For example, the electronic device (101) may identify (or obtain) text information of a translated text object within a translated image to be generated using text information of the image (or a text object within the image). As a non-limiting example, the electronic device (101) may identify (or obtain) placement information indicating a position of a translated text object within a translated image to be generated using placement information of the image (or a text object within the image).

[0159] In one example, the electronic device (101) can perform translation of first text objects of a first source language among text objects within an image substantially simultaneously based on the AI ​​model (340). After performing translation of the first text objects, the electronic device (101) can perform translation of second text objects of a second source language among the text objects within the image substantially simultaneously based on the AI ​​model (340). In other words, the electronic device (101) can perform translation based on the AI ​​model (340) for text objects having the same source language among the text objects within the image.

[0160] For example, the electronic device (101) may generate text objects translated into the target language and then generate a translated image including the translated text objects. In one example, the electronic device (101) may adjust the properties of the translated text objects to correspond to the properties of the image based on the AI ​​model (340). For example, the electronic device (101) may identify the properties of the image based on the AI ​​model (340). For example, the properties of the image may include the location of the image, the location and color of elements (e.g., images) included in the image, the representative color of the image, or the size of the image. For example, the electronic device (101) may adjust the properties of the translated text objects according to the representative color of the image. For example, the properties of the text objects may include at least one of the color, thickness, size, font, font size, or background color of the text object.

[0161] As a non-limiting example, the electronic device (101) may generate a translated image based on replacing text objects in an image before translation with translated text objects. For example, the electronic device (101) may generate the translated image based on removing text objects in the image before translation (or outpainting) and adding translated texts in the image before translation (or inpainting). Alternatively, for example, the electronic device (101) may generate the translated image based on adding translated texts above the text objects in the image before translation (or inpainting). In this case, the text objects before translation may be visually invisible due to the background color of the translated texts in the translated image. The background color may be adjusted according to a representative color of the image before translation. In other words, the electronic device (101) can generate the translated image by inpainting the translated texts for some of the areas where the text objects before translation of the image before translation are located, and inpainting the representative color of the first image for other parts of the areas. As a non-limiting example, the electronic device (101) can generate (or obtain) the translated image in which the text objects have been replaced using the AI ​​model (340) by inputting the text information of the image before translation and the text objects in the image before translation into the AI ​​model (340).

[0162] In the above example, the electronic device (101) is described as performing translation on the image based on the AI ​​model (340), but the present disclosure is not limited thereto. As a non-limiting example, the processor (310) of the electronic device (101) (or a program executed by the processor (310)) can perform translation on the image.

[0163] In operation (640), the electronic device (101) may store the translated image within a resource map. For example, the resource map may be stored within a storage area utilized by the application. For example, the electronic device (101) may store the resource map for one or more elements of the web page within the storage area as the application for the web browser is executed. For example, the resource map may include image information of an element having an image property displayed within the web page (or information on pixels constituting an image). For example, the resource map may include an image for which translation is performed in operation (635). For example, the resource map may include the image including a text object before being translated.

[0164] For example, the electronic device (101) may store the translated image within the resource map. For example, the electronic device (101) may store both the image before translation and the translated image within the resource map.

[0165] In operation (645), the electronic device (101) may display a UI including a translated image. For example, the electronic device (101) may display a UI (or representation) including the translated image changed from the image before translation on the display (320) by identifying the translated image stored in the resource map. As a non-limiting example, the electronic device (101) may display a first representation of the web page including the image before translation, and then, after performing a translation on the image, display a second representation of the web page including the translated image. In this case, the second representation may include an image (or elements) in which an image (or elements) in the first representation is translated. The electronic device (101) may generate and store the translated image based on identifying the event, and then, by identifying the translated image instead of the image before translation in the resource map, display the UI including the translated image replaced from the image before translation.

[0166] In the following FIGS. 7a to 9, specific examples of a method for performing translation for an element (or image) having an image property and an element (or text) having a text property based on an AI model (340), as described in FIGS. 5 and 6, are described.

[0167] Figure 7a illustrates an example of a method for requesting translation of texts and images among elements of a web page based on an AI model.

[0168] FIG. 7A illustrates an example (700) of a method in which an electronic device (101) requests translation of elements of a web page based on an AI model (340) to perform translation of the web page. For example, the method of FIG. 7A may exemplify specific operations for operations (605) and (610) of FIG. 6 .

[0169] Referring to example (700), the electronic device (101) can display a UI (710) on the display (320). For example, while displaying the UI (710) representing a web page of the web browser, the electronic device (101) can identify an event for translating the web page. For example, by identifying the event for translating the web page, the electronic device (101) can identify configuration information of the web page.

[0170] For example, the electronic device (101) can identify elements within the UI (710), which is a portion currently being displayed on the display (320), among the web pages using the configuration information. For example, the area corresponding to the UI (710) among the web pages can be referred to as a first area or a viewport. For example, the electronic device (101) can identify images (721, 722, 723, 724), which are elements having image properties within the web pages, using the configuration information. For example, the electronic device (101) can identify texts (731, 732, 733), which are elements having text properties within the web pages, using the configuration information. For example, the texts (731, 732, 733) can be texts that are distinct from the images (721, 722, 723, 724). In other words, the texts (731, 732, 733) may be elements with text properties, rather than text within the images (721, 722, 723, 724).

[0171] In the above example, the configuration information is used to identify elements within the UI (710), which is the currently displayed portion, but the present disclosure is not limited thereto. For example, the electronic device (101) may also use the configuration information to identify elements within the UI (710), which is the currently displayed portion, and within a portion of the web page that extends from the UI (710) and is currently undisplayed. For example, the portion of the web page that is currently undisplayed may be referred to as a second region. For specific details related thereto, reference may be made to FIG. 7B below.

[0172] For example, the electronic device (101) may request a translation based on the AI ​​model (340) for the texts (731, 732, 733) and images (721, 722, 723, 724) identified using the configuration information. For example, the electronic device (101) may transmit the texts (731, 732, 733) and images (721, 722, 723, 724) to the AI ​​model (340) according to priority.

[0173] For example, the priority may be determined by considering the translation speed (or processing speed) based on the AI ​​model (340). For example, the priority for text translation may be higher than the priority for image translation. For example, among images, the priority of an image whose size exceeds a reference size may be higher than the priority of an image whose size is smaller than the reference size. For example, the electronic device (101) may select an image whose size exceeds the reference size as a target for translation among the images. An image whose size is smaller than the reference size may be selected as a target for translation after an image whose size exceeds the reference size is translated. For example, the reference size may be determined as a specified magnification (e.g., 1 / 5) with respect to the size of the UI (710) currently being displayed. For example, when the number of elements requesting translation exceeds the reference number, translation may be requested first for elements corresponding to the reference number, and then translation may be requested for the remaining elements. For example, the above-mentioned reference number may be determined by considering the translation speed (or processing speed) based on the AI ​​model (340). In the present disclosure, for convenience of explanation, it is assumed that the above-mentioned reference number is 3.

[0174] With reference to example (700) in relation to the above priorities, the electronic device (101) may perform translation of texts (731, 732, 733) before images (721, 722, 723, 724). For example, the electronic device (101) may transmit a request (741) (or a list, a translation list) for texts (731, 732, 733) to the AI ​​model (340) and then request translation of images (721, 722, 723, 724). At this time, the electronic device (101) may transmit a request (742) (or a list, a translation list) for an image (722) exceeding the reference size among the images (721, 722, 723, 724) and an image (721, 723) smaller than the reference size to the AI ​​model (340). Although not shown in the example (700), the electronic device (101) may transmit a request (742) including elements less than the reference number to the AI ​​model (340), and then transmit a request including the image (724), which is the remaining element, to the AI ​​model (340). In the example, the images (721, 722, 723) may be included in a first image set corresponding to the reference number, and the image (724) may be included in a second image set.

[0175] For example, the AI ​​model (340) may perform translation according to the order of received requests (741, 742). In one example, while the translation (or process for translation) of the texts (731, 732, 733) of the request (741) is being performed, the translation (or process for translation) of the images (721, 722, 723) of the request (742) may be performed in parallel. In relation to the process for translation, for example, in the case of a request (742) including images among the requests (741, 742), the AI ​​model (340) may determine whether the image includes text, and if the image includes text, extract text information. For specific details related thereto, reference may be made to FIGS. 8A and 8B below.

[0176] FIG. 7b illustrates an example of a method for requesting translation of images in a first area of ​​a web page and images in a second area of ​​a web page based on an AI model.

[0177] FIG. 7B illustrates an example (750) of a method in which an electronic device (101) requests translation of images (721, 722, 723, 724) within a first area of ​​a web page and images (771, 772, 773) within a second area of ​​the web page to perform translation of the web page based on an AI model (340). For example, the method of FIG. 7A may exemplify specific operations for operations (605) and (610) of FIG. 6 .

[0178] In the example (750) of FIG. 7B, for convenience of explanation, the electronic device (101) requests translation of images (721, 722, 723, 724) within the UI (710) corresponding to the first area of ​​the web page and requests translation of images (771, 772, 773) within the second area (760), but the present disclosure is not limited thereto. For example, the electronic device (101) may request translation of images (721, 722, 723, 724) within the UI (710) corresponding to the first area and then request translation of texts within the second area (760). The electronic device (101) may request translation of images (771, 772, 773) within the second area (760) after requesting translation of texts within the second area (760).

[0179] In the example (750), the second area (760) is illustrated as an area extending in a first direction (e.g., from top to bottom) from the UI (710), which is the first area (or viewport), but the present disclosure is not limited thereto. For example, the second area (760) may include an area extending in a second direction (e.g., from bottom to top) opposite to the first direction from the UI (710), which is the first area. As a non-limiting example, the size of the second area (760) may be set to a specified magnification (e.g., 2 times) with respect to the size of the UI (710). In one example, the second area (760) may correspond to the first area and may be displayed on the display (320) according to an input (or user input) received with respect to the UI (710) displayed on the display (320). As a non-limiting example, the input may include a scroll input regarding the UI (710).

[0180] Referring to example (750), the electronic device (101) may transmit a request (742) to the AI ​​model (340) for images (722) exceeding the reference size and images (721, 723) smaller than the reference size among the images (721, 722, 723, 724) in the UI (710) corresponding to the first area. Among the images (721, 722, 723, 724) of the first area, the image (724) may be excluded from the request (742) according to the reference number. After transmitting the request (742) to the AI ​​model (340), the electronic device (101) may transmit a request (743) (or a list, a translation list) including the image (724) of the first area and the images (771, 772) of the second area (760) to the AI ​​model (340). In the example, the image (724) and the images (771, 772) may be included in the second image set. Among the images (771, 772, 773) of the second area (760), the image (773) may be excluded from the request (743) according to the reference number. Although not illustrated in FIG. 7B, the electronic device (101) may request a translation of the image (773) after transmitting the request (743).

[0181] In the examples of FIGS. 7A and 7B, a case is illustrated where a translation is requested first for an image exceeding the reference size among the images, but the present disclosure is not limited thereto. For example, the electronic device (101) may transmit a request for translation for an image to the AI ​​model (340) according to the resolution of the images. For example, the electronic device (101) may perform translation for an image exceeding a reference resolution (e.g., 512x512) among elements (or images) within a web page based on the AI ​​model (340). In addition, in one example, the electronic device (101) may divide an image exceeding another reference size exceeding the reference size into a plurality of partial images and then perform translation on the divided partial images. For example, the partial images may be divided based on the location of text within the image exceeding the other reference size.

[0182] FIG. 8a illustrates an example of a method for determining whether an image on a web page contains a text object based on an AI model.

[0183] FIG. 8A illustrates an example (800) of a method by which an electronic device (101) determines whether an image of a web page contains a text object based on an AI model (340). For example, the method of FIG. 8A may exemplify specific operations for operations (610) and (620) of FIG. 6. The UI (710) of FIG. 8A may be substantially identical to the UI (710) of FIG. 7A.

[0184] For example, the electronic device (101) may use the configuration information of the web page to identify images (721, 722, 723, 724), which are elements having image properties among the elements of the web page, and transmit a request for translation (e.g., request (742) of FIG. 7A) for at least some of the images (721, 722, 723, 724) (e.g., images (721, 722, 723)) to the AI ​​model (340). In the example (800) of FIG. 8A, it is assumed that the request (742) is transmitted to the AI ​​model (340). However, the present disclosure is not limited thereto.

[0185] For example, the electronic device (101) can determine whether an image input to the AI ​​model (340) is an image having a text object. For example, the electronic device (101) can determine whether the image (721) includes a text object based on the AI ​​model (340) (or the first AI engine of the AI ​​model (340). For example, the electronic device (101) can generate an output (831) indicating whether the image (721) includes a text object based on the AI ​​model (340). For example, the electronic device (101) can determine whether the image (722) includes a text object based on the AI ​​model (340). For example, the electronic device (101) can generate an output (832) indicating whether the image (722) includes a text object based on the AI ​​model (340). For example, the electronic device (101) may determine whether the image (723) includes a text object based on the AI ​​model (340) (or the first AI engine of the AI ​​model (340)). For example, the electronic device (101) may generate an output (833) indicating whether the image (723) includes a text object based on the AI ​​model (340).

[0186] For example, the output (831) may indicate a first value (e.g., 1) (or true) indicating that the image (721) contains a text object (or that the image (721) is an image that contains a text object). For example, the output (832) may indicate the first value indicating that the image (722) contains a text object. For example, the output (833) may indicate a second value (e.g., 0) (or false) indicating that the image (723) does not contain a text object (or that the image (721) is an image that does not contain a text object).

[0187] In example (800), since image (721) includes a text object ('with a puppy') and image (722) includes text objects ('A shopping', 'today's new arrivals'), output (831) and output (832) can each indicate the first value. Furthermore, in example (800), since image (723) does not include a text object, output (833) can indicate the second value.

[0188] Fig. 8b illustrates an example of a method for obtaining text information of a text object included in an image when the image of a web page includes a text object, based on an AI model.

[0189] FIG. 8B illustrates an example (850) of a method by which an electronic device (101) acquires (or extracts) text information of a text object included in an image of a web page based on an AI model (340). For example, the method of FIG. 8B may exemplify specific operations for the operation (630) of FIG. 6. The UI (710) of FIG. 8B may be substantially identical to the UI (710) of FIG. 7A.

[0190] Referring to FIG. 8B, after determining that each of the images (721, 722) includes a text object, the electronic device (101) may obtain text information of the text object included in each of the images (721, 722) based on the AI ​​model (340). For example, the electronic device (101) may obtain text information (860) of the text object included in the image (721) based on the AI ​​model (340) (or the first AI engine of the AI ​​model (340). For example, the electronic device (101) may obtain text information (870) of the text object included in the image (722) based on the AI ​​model (340) (or the first AI engine of the AI ​​model (340).

[0191] For example, the text information (860) may include an OCR result (861a) and position information (861b) (or placement information) for the text object ('with a dog') of the image (721). For example, the position information (861b) may include coordinates (e.g., x, y) at which the text object ('with a dog') is located within the image (721) (or coordinates at which a virtual block (865) (or text box) is located), the size of the block (865) (e.g., w, h), and rotation information (e.g., angle) indicating the degree to which the block (865) is rotated. For example, the coordinates (e.g., x, y) may be defined from a reference position of the image (721) (e.g., a point at the upper left or the center of the image (721)). For example, among the sizes, w may represent the horizontal length of the block (865), and h may represent the vertical length of the block (865). For example, the rotation information (e.g., angle) may represent the degree to which the block (865) is rotated with respect to the image (721) (or, in a specified direction (e.g., horizontal direction) of the image (721)).

[0192] For example, the text information (870) may include an OCR result (871a) and location information (871b) for a text object ('A Shopping') in the image (722). For example, the location information (871b) may include coordinates (e.g., x, y) at which the text object ('A Shopping') is located within the image (722) (or coordinates at which a virtual block (875) (or text box) is located), a size (e.g., w, h) of the block (875), and rotation information (e.g., angle) indicating a degree to which the block (875) is rotated. In addition, for example, the text information (870) may include an OCR result (872a) and location information (872b) at which the text object ('Today's New Arrivals') in the image (722). For example, the location information (872b) may include coordinates (e.g., x, y) where the text object ('Today's New Arrivals') is located within the image (722) (or coordinates where the virtual block (876) is located), the size of the block (876) (e.g., w, h), and rotation information (e.g., angle) indicating the degree to which the block (876) is rotated.

[0193] For example, the electronic device (101) can identify the language (or source language) of a text object in an image based on the AI ​​model (340) (or the first AI engine). For example, the electronic device (101) can identify that the language of the text object ('with a puppy') in the image (721) is Korean, identify that the language of the text object ('A shopping') in the image (722) is Korean and English, and identify that the language of the text object ('today's shopping') in the image (722) is Korean based on the AI ​​model (340).

[0194] In the above example, the text information is described as including OCR results and location information for a text object included in an image, but the present disclosure is not limited thereto. For example, the electronic device (101) may obtain text information further including text attributes for a text object included in an image based on the AI ​​model (340) (or the first AI engine of the AI ​​model (340). For example, the text attributes may include at least one of the color, thickness, size, font, font size, or background color of the text object.

[0195] Figure 8c illustrates an example of a method for generating a translated image using image and text information based on an AI model.

[0196] FIG. 8C illustrates an example (880) of a method in which an electronic device (101) generates a translated image using an image of a web page and text information of a text object included in the image based on an AI model (340). For example, the method of FIG. 8C may exemplify specific operations for operation (635) of FIG. 6. The method of example (880) assumes a case in which a translated image (890) is generated using the image (722) of example (850) of FIG. 8B.

[0197] Referring to example (880), the electronic device (101) can generate a translated image (890) based on the AI ​​model (340) (or the second AI engine of the AI ​​model (340)) using the image (722) and the text information (870) about the text objects of the image (722). For example, the electronic device (101) can obtain text objects ('A shopping', 'Today's Arrivals') translated into a target language (e.g., English) based on the AI ​​model (340) using the image (722) and the text information (870). For example, the electronic device (101) can generate a translated image (890) including the translated text objects ('A shopping', 'Today's Arrivals') based on the AI ​​model (340).

[0198] In the example (880) of FIG. 8c, an image (890) is generated that includes text objects ('A shopping' and 'Today's Arrivals') that include one source language (e.g., Korean) in the image (722), translated into one target language (e.g., English) text objects ('A shopping', 'Today's Arrivals'), but the present disclosure is not limited thereto. For example, even when an image includes multiple text objects, and the multiple text objects include multiple languages, the electronic device (101) may generate a translated image that includes text objects translated from the multiple text objects of the multiple languages ​​based on the AI ​​model (340).

[0199] In one example, assume that an image includes a first text object and a second text object in a first language (e.g., Korean) different from the target language (e.g., English), and that the image includes a third text object and a fourth text object in a second language (e.g., French) different from the target language. For example, when performing a translation of the image, the electronic device (101) may input the first text object and the second text object in the first language into an AI model (340), thereby obtaining a fifth text object and a sixth text object, respectively, in which the first text object and the second text object of the image are translated into the target language. After performing the translation of the first text object and the second text object, the electronic device (101) may input the third text object and the fourth text object in the second language into an AI model (340), thereby obtaining a seventh text object and an eighth text object, respectively, in which the third text object and the fourth text object of the image are translated into the target language. The electronic device (101) can generate a translated image including the fifth text object, the sixth text object, the seventh text object, and the eighth text object from the image including the first text object, the second text object, the third text object, and the fourth text object.

[0200] Also, in one example, the electronic device (101) is illustrated as generating an image (890) that includes text objects ('A shopping' and 'Today's Arrivals') that include one source language (e.g., Korean) in the image (722) translated into one target language (e.g., English), but the present disclosure is not limited thereto. For example, the electronic device (101) may also generate a plurality of translated images that include text objects ('A shopping' and 'Today's Arrivals') that include one source language (e.g., Korean) in the image (722) translated into a plurality of target languages ​​(e.g., English, Spanish).

[0201] In addition, in one example, when generating an image (890) translated from an image (722), the electronic device (101) may further generate summary information of the image (722) based on the AI ​​model (340). For example, the electronic device (101) may generate the summary information (e.g., shopping information about a new item from A Shopping) describing the image (722) based on the AI ​​model (340). For example, the electronic device (101) may display the translated image (890) that has been changed (or replaced) from the image (722) on the UI (710). At this time, the electronic device (101) may also display a visual object (or the summary information) representing the summary information of the translated image (890). In one example, the electronic device (101) may also display the summary information according to a user input identified with respect to the translated image (890).

[0202] In addition, in one example, when generating a translated image (890) from an image (722), the electronic device (101) may adjust the properties of the translated text objects based on the AI ​​model (340). For example, the electronic device (101) may adjust the location and properties of the text objects ('A shopping', 'Today's Arrivals') in the image (722) based on the AI ​​model (340) (or the second AI engine of the AI ​​model (340). For example, the electronic device (101) may adjust the properties of the text objects ('A shopping', 'Today's Arrivals') according to the representative colors of the pixels of the image (722) based on the AI ​​model (340). For example, the properties may include at least one of the color, thickness, size, font, font size of the text, or the background color of the text.

[0203] Additionally, in one example, the electronic device (101) may obtain and store additional information in addition to the translated image (890) from the image (722). For example, the additional information may include at least one of an image representing a portion of the translated image (890) including translated text objects ('A shopping', 'Today's Arrivals'), rather than the entire translated image (890), text objects before translation ('A shopping' and 'Today's New Arrivals'), or text objects after translation ('A shopping', 'Today's Arrivals').

[0204] As described above, the electronic device (101) can generate a translated image (890) (and additional information) and then store the translated image (890) within a resource map. For specific details on how to display a UI using the image stored within the resource map, reference may be made to FIG. 9 below.

[0205] Figure 9 illustrates an example of a method for displaying a web page using a resource map that stores images and translated images.

[0206] FIG. 9 illustrates an example of a method for displaying a web page using a resource map (950) that stores an image and an image translated based on an AI model (340) using the image.

[0207] Referring to FIG. 9, the electronic device (101) can display a UI (710) on the display (320). For example, the UI (710) can represent a web page of the web browser. For example, the UI (710) can include images (721, 722) and texts (731, 732, 733).

[0208] For example, the electronic device (101) may store a resource map (950) for one or more elements (e.g., images (721, 722)) of the web page in a storage area of ​​the electronic device (101) when executing an application for the web browser. For example, the resource map (950) may include image information (or information on pixels constituting an image) of an element having an image property displayed in the web page. For example, the resource map (950) may store the image (721) in a portion (951) related to the image (721). For example, the resource map (950) may store the image (722) in a portion (952) related to the image (722). However, the present disclosure is not limited thereto.

[0209] For example, the electronic device (101) may identify an event for translation of the web page while displaying the UI (710). For example, the electronic device (101) may perform translation of the UI (710) upon identifying the event for translation of the web page. For example, the electronic device (101) may use images (721, 722) and texts (731, 732, 733) in the UI (710) to generate translated images (921, 922) and translated texts (931, 932, 933) based on the AI ​​model (340). For example, the electronic device (101) may store the translated images (921, 922) in the resource map (950). For example, the translated image (921) may be stored in the portion (951). For example, the translated image (922) may be stored within the portion (952). For example, the translated image (922) may be an example of the translated image (890) of FIG. 8C.

[0210] For example, the electronic device (101) can display the UI (910) on the display (320) by identifying stored images (921, 922) in the resource map (950). For example, the stored images (921, 922) can be referenced as cached images. For example, the translated image (921) can be changed from (or replaced with) the image (721). For example, the translated image (922) can be changed from (or replaced with) the image (722).

[0211] For example, the electronic device (101) may display a UI (910) including texts to be translated (931, 932, 933) that have been changed from texts (731, 732, 733) on the display (320). For example, the translated text (931) may be changed from (or replaced by) the text (731). For example, the translated text (932) may be changed from (or replaced by) the text (732). For example, the translated text (933) may be changed from (or replaced by) the text (733).

[0212] In one example, the electronic device (101) may identify another event (or another input) for displaying an image of the web page before translation. For example, the electronic device (101) may identify another event including an input for an icon representing a translation function included in a menu of a UI (e.g., UI (710), UI (910)). For example, upon identifying the other event, the electronic device (101) may display a UI (710) modified from the UI (910). For example, the electronic device (101) may display the UI (710) by identifying images (721, 722) stored in the resource map (950). Thereafter, when the electronic device (101) again identifies an input for an icon representing a translation function included in the menu, the electronic device (101) may again display the UI (910) modified from the UI (710). The electronic device (101) can display the UI (910) by identifying the images (921, 922) stored in the resource map (950). At this time, the electronic device (101) can use the images (921, 922) stored in the resource map (950) rather than regenerating the images (921, 922) based on the AI ​​model (340) using the images (721, 722). Accordingly, the UI (910), which is a translated web page, can be displayed more quickly than regenerating the images (921, 922) based on the AI ​​model (340) using the images (721, 722).

[0213] In the above example, the electronic device (101) is described as displaying a UI (710) before translation and a UI (910) after translation according to an input for the icon, but the present disclosure is not limited thereto. For example, the electronic device (101) may display a UI (910) including an image (722) that has been changed (or toggled) from an image (922) by obtaining an input with respect to an image (922) within the UI (910). In this case, the UI (910) may include an image (921) and an image (722).

[0214] In the example of FIG. 9, the electronic device (101) is illustrated as displaying a translated UI (910) by identifying an input (or an event) while displaying the web page, but the present disclosure is not limited thereto. For example, the electronic device (101) may display another web page changed from the web page on the display (320) by identifying an input to the UI (910) of the web page. The resource map (950) may be maintained (or managed, operated) according to the lifecycle of the web page. For example, the lifecycle of the web page may be related to the execution of the web page (e.g., execution in the foreground or execution in the background / foreground) or the display of the web page. For example, when displaying the other web page, the electronic device (101) may reset (or delete, remove) the resource map (950) acquired and stored in relation to the web page. Thereafter, the electronic device (101) can display the web page again by identifying the input for the other web page. Accordingly, the electronic device (101) can re-acquire the resource map (950) related to the web page. At this time, the re-acquired resource map (950) may include images (721, 722) and may not include images (921, 922). Alternatively, for example, the electronic device (101) may suspend (or terminate) the execution of the application providing the web browser. Accordingly, the electronic device (101) can reset (or delete, remove) the resource map (950) acquired and stored related to the web page. In one example, when the resource map (950) is reset, the electronic device (101) may store (or maintain storage of) the translated images (921, 922) within the storage area (or in an application that stores images, e.g., a gallery).As a non-limiting example, from the time the resource map (950) is reset, the electronic device (101) may store the translated images (921, 922) within the storage area for a predetermined period of time. For example, the length of the predetermined period of time may be preset by the user or determined based on recorded information about the web page (e.g., average usage time or usage pattern of the web page). As a non-limiting example, when the resource map (950) is reset, the electronic device (101) may store a file including the translated images (921, 922) within the storage area. Alternatively, for example, when the resource map (950) is reset, the electronic device (101) may store the translated images (921, 922) by transmitting the translated images (921, 922) to an external electronic device (e.g., a server).

[0215] In the above example, a case is described where the resource map (950) is reset in response to a change in the web page, but the present disclosure is not limited thereto. For example, upon identifying an input for displaying another web page from the web page, the electronic device (101) may display a notification object on the display (320) that inquires whether to perform a reset of the resource map (950) related to the web page. For example, upon obtaining an input for the visual object, the electronic device (101) may refrain from resetting the resource map (950) even when displaying the other web page. Alternatively, for example, the electronic device (101) may, upon identifying an input for displaying the other web page from the web page, display a notification object on the display (320) that inquires whether to store translated images (921, 922) stored in the resource map (950) related to the web page in another storage area of ​​the electronic device (101) distinct from the storage area that stores the resource map (950). Upon receiving an input for the notification object inquiring whether to store translated images (921, 922) in the resource map (950) in the other storage area, the electronic device (101) may store the translated images (921, 922) in the resource map (950) in the other storage area and delete the resource map (950) from the storage area.

[0216] In addition, for example, the electronic device (101) may refrain from resetting the resource map (950) even when displaying the other web page, upon identifying an input for displaying the other web page from the web page, according to the following conditions. For example, the conditions may include a case where an application using the resource map (950) is running in the foreground. Or, for example, the conditions may include a case where a transition to the other web page is made through a link (or hyperlink) of the web page running in the same tab (or menu) of the web browser (a transition to a related web page). In other words, a reset of the resource map (950) may be performed when the other web page is run in a different tab of the web browser or a transition is made by inputting a URL indicating the other web page. Or, for example, the conditions may include a case where a certain amount of time (e.g., 5 minutes) has elapsed since the transition from the web page to the other web page. If the resource map (950) is not reset according to the above conditions, the electronic device (101) can relatively quickly provide the translated UI of the web page when switching back to the web page from the other web page.

[0217] In one example, the electronic device (101) may delete images (921, 922) stored in the resource map (950) when changing from the web page to another web page. At this time, the electronic device (101) may store translated texts (or text information (e.g., OCR results)) of the images (921, 922) in the storage area (or the electronic device (101)) that is distinct from the resource map (950). For example, the electronic device (101) may re-generate images (921, 922) by using the translated texts (or text information) stored in the electronic device (101) when a translation of the web page is requested again. For example, when a translation of the web page is requested again, the electronic device (101) may skip determining whether the images (721, 722) contain text objects and regenerate the translated images (921, 922) by using the text information stored within the electronic device (101).

[0218] In one example, the electronic device (101) may store at least one of the translated text objects or the pre-translation text objects obtained when generating the translated images (921, 922) in a clipboard running on the electronic device (101). As a non-limiting example, the electronic device (101) may display (or input) the translated text objects and / or the pre-translation text objects stored in the clipboard on the display (320) based on receiving an input.

[0219] In one example, the electronic device (101) can avoid resetting the resource map (950) by identifying historical information about the web page. For example, the historical information may include the frequency of use or number of accesses to the web page. For example, the electronic device (101) can store translated images (921, 922) of the resource map (950) within the electronic device (101) for a certain period of time based on the historical information.

[0220] In one example, the electronic device (101) may not only display (or provide) the UI (910) of the translated web page on the display (320), but may also provide other types of output information. For example, the other types of information may include auditory information. For example, the electronic device (101) may generate audio (or voice) for a translated text object of a translated image (922) and provide the generated audio through an output device (e.g., the audio output module (155) of FIG. 1). For example, the electronic device (101) may provide audio for a text object in an image (722) or audio for a text object in a translated image (922), depending on a user's input. Furthermore, in one example, the electronic device (101) may also provide the result of the translation of the audio together in the UI (910).

[0221] In FIG. 9, translated images (921, 922) are generated for static images (721, 722) within the UI (710), but the present disclosure is not limited thereto. For example, if the UI (710) includes a dynamic image (or a video, a graphics interchange format (GIF)), text objects within the dynamic image may also be translated based on the AI ​​model (340). For example, the electronic device (101) may translate a text object within a representative image (or a thumbnail, a specific frame) of the dynamic image and display the translated image within the UI (910). For example, the representative image may be set within the dynamic image or selected based on the AI ​​model (340). Alternatively, for example, the UI (710) may include a captured screen for screen information of a software application. For example, the captured screen may be obtained through screen reading (or screen capture), and translation may be performed on the captured screen.

[0222] Figure 10 illustrates an example of a method for generating a translated image using another web page related to a web page based on an AI model and displaying a web page including the translated image.

[0223] FIG. 10 illustrates an example (1000) of a method for displaying a UI (1030) of a web page including a translated image using another web page related to the web page, based on an AI model (340). The method of FIG. 10 can be performed by the electronic device (101) of FIG. 3b.

[0224] Referring to example (1000), the electronic device (101) can display the UI (1010) of the web page on the display (320). For example, the UI (1010) can include an image (1011). For example, the image (1011) can include a text object (1013). For example, the text object (1013) can include 'Your Health Partner' entered in a source language (e.g., Korean). In this case, the web page can be a web page related to the source language (e.g., a Korean server of the web page). For example, the electronic device (101) can perform a translation of the UI (1010) based on the AI ​​model (340). For example, the electronic device (101) can perform a translation of the UI (1010) by identifying an event for the translation. The above translation assumes translation into a target language (e.g. English).

[0225] For example, in response to identifying the event, the electronic device (101) may identify (or retrieve) the other web page. For example, the other web page may be a web page (e.g., a US server of the web page) associated with a target language (e.g., English). For example, the electronic device (101) may identify the UI (1020) of the other web page. For example, the UI (1020) may include an image (1021). For example, the image (1021) may include a text object (1023). For example, the text object (1023) may include 'Your Wellness Partner' entered in the target language (e.g., English). For example, the electronic device (101) may identify (or obtain) configuration information of the UI (1020).

[0226] For example, the electronic device (101) may provide configuration information of the UI (1020) to the AI ​​model (340) to perform the above translation for the UI (1010). For example, the electronic device (101) may generate the UI (1030) by providing the configuration information of the UI (1020) to the AI ​​model (340). For example, the UI (1030) may include a translated image (1031). For example, the image (1031) may include a translated text object (1033). For example, the text object (1033) may include 'Your Wellness Partner' entered in a target language (e.g., English).

[0227] In the example (1000) of FIG. 10, for convenience of explanation, the image (1011) of the web page and the image (1021) of the other web page are depicted as being identical, but the present disclosure is not limited thereto. For example, the image (1011) of the web page may be different from (or similar to) the image (1021) of the other web page.

[0228] Figures 11a and 11b illustrate examples of web pages provided upon activation of the translation function.

[0229] FIGS. 11A and 11B illustrate examples (1101, 1102, 1103, 1104, 1105, 1106, 1107, 1108) of a UI (or representation) of a web page displayed through a display (320) provided by an electronic device (101) upon activation of a translation function.

[0230] Referring to example (1101), the electronic device (101) may display a UI (1110) through the display (320). For example, the UI (1110) may represent at least a portion of the web page. For example, the at least portion may include an image (1113). In addition, for example, the UI (1110) may include a menu (1115) for executing a function for the web page. For example, the menu (1115) may include a plurality of icons representing the function. For example, the plurality of icons may include an icon (1117) including a translation function. For example, the icon (1117) may include a function performed based on the AI ​​model (340). In example (1101), the image (1113) may include text objects in a first language (e.g., Korean).

[0231] In the UI (1110) of Example (1101), the electronic device (101) can receive (or identify) an input for an icon (1117). Referring to Example (1102), the electronic device (101) can display an execution menu (1120) popping up on the icon (1117) based on receiving an input for the icon (1117). As a non-limiting example, the execution menu (1120) can include a first execution menu (1121) and a second execution menu (1122). For example, the first execution menu (1121) can indicate a summary function for the web page. For example, the second execution menu (1122) can indicate a translation function for the web page.

[0232] In the UI (1110) of Example (1102), the electronic device (101) may receive (or identify) an input for a second execution menu (1122). For example, based on receiving an input for the second execution menu (1122), the electronic device (101) may perform a translation for the web page. Performing the translation for the web page may refer to Example (1103).

[0233] In example (1103), the electronic device (101) may perform a translation on the web page and display the translated UI (1110) of the web page (or the UI (1110) changed from the UI (1110) of example (1101)) through the display (320). As a non-limiting example, as the translation is performed, the electronic device (101) may provide a visual effect (1139) that progresses from the top to the bottom of the UI (1110) of example (1103). For example, the visual effect (1139) may include a gradient effect to indicate that the translation is being performed. As the visual effect (1139) progresses from the top to the bottom of the UI (1110) of example (1103), elements of the web page located within the visual effect (1139) may be changed (or replaced) with translated elements. For example, translated texts (1135) of the web page may be displayed within the UI (1110) of example (1103). At this time, images (1113) positioned outside of the visual effect (1139) may remain untranslated. However, the present disclosure is not limited thereto. For example, the visual effect (1139) may be used only as a visual effect to indicate to the user that translation is being performed, and elements of the web page positioned within the visual effect (1139) may not be changed to translated elements. After the visual effect (1139) starts from the top of the UI (1110) and progresses completely to the bottom, and after the display of the visual effect (1139) is stopped, the elements of the web page may be changed to translated elements simultaneously (or, all at once, immediately, sequentially). As a non-limiting example, the electronic device (101) may display a translation menu (1130) within the UI (1110) of example (1103).For example, a translation menu (1130) may be displayed within the UI (1110) of example (1103) based on receiving input for the second execution menu (1122). As a non-limiting example, the translation menu (1130) may include an icon (1131) indicating a currently selected target language (e.g., English) and an icon (1132) indicating that a translation is being performed. For example, within the UI (1110) of example (1103), while a visual effect (1139) is displayed, the icon (1132) may include an image indicating that the translation is being performed.

[0234] Referring to example (1104), the electronic device (101) may display a translated UI (1110) of the web page (or a UI (1110) changed from the UI (1110) of example (1101)) through the display (320) as the translation of the web page is completed. For example, the electronic device (101) may display translated elements of the web page within the UI (1110) of example (1104). For example, the UI (1110) of example (1104) may include a translated image (1143) including translated texts (1135) and translated text objects. For example, the translated image (1143) may be an image in which the text objects of the image (1113) are changed (or replaced) with translated text objects.

[0235] As a non-limiting example, the electronic device (101) may display a translation menu (1130) within the UI (1110) of example (1104). For example, as translation is completed, the translation menu (1130) may include an icon (1131) indicating the currently selected target language and an icon (1142) indicating that translation is complete. For example, within the UI (1110) of example (1104), as translation is completed, the display of the visual effect (1139) may stop, and the icon (1142) may include texts ('View Original') to provide a display of the web page before translation.

[0236] The electronic device (101) may receive an input for an icon (1131) indicating a currently selected target language of a translation menu (1130). Referring to example (1105), upon receiving (or identifying) an input for the icon (1131), the electronic device (101) may display a third execution menu (1150) indicating selectable target languages. For example, the third execution menu (1150) may include a menu (1151) indicating a currently selected target language (e.g., English), a menu (1152) indicating selectable target languages ​​(e.g., Japanese), and a menu (1153) for adding selectable target languages. For example, upon receiving an input for the menu (1153), the electronic device (101) may add a selectable target language.

[0237] The electronic device (101) may perform a translation of the web page upon receiving an input for the menu (1152). Referring to example (1106), the electronic device (101) may perform a translation into a changed target language (e.g., Japanese) upon receiving an input for the menu (1152). As a non-limiting example, when the translation into the changed target language is performed, the electronic device (101) may display the original text of the web page and then perform a translation into the changed target language. For example, the electronic device (101) may re-display the UI (1110) of example (1101) that has been changed from the UI (1110) of example (1104) and perform a translation into the changed target language. However, the present disclosure is not limited thereto. For example, the translation into the changed target language may be performed while the UI (1110) of example (1104) is displayed.

[0238] In example (1106), the electronic device (101) may perform a translation of the web page and display the translated UI (1110) of the web page (or the UI (1110) changed from the UI (1110) of example (1101)) through the display (320). As a non-limiting example, as the translation into the changed target language is performed, the electronic device (101) may provide a visual effect (1169) that progresses from the top to the bottom of the UI (1110) of example (1106). The specific details of the visual effect (1169) may be substantially identical to the details of the visual effect (1139). For example, as the visual effect (1169) progresses from the top to the bottom of the UI (1110) of example (1106), elements of the web page located within the visual effect (1169) may be changed (or replaced) with translated elements. For example, translated texts (1165) of the above web page may be displayed within the UI (1110) of example (1106). In this case, images (1113) positioned outside of the visual effects (1169) may remain untranslated.

[0239] As a non-limiting example, the electronic device (101) may display a translation menu (1130) within the UI (1110) of example (1106). For example, the translation menu (1130) may be displayed within the UI (1110) of example (1106) based on receiving an input for the second execution menu (1122). As a non-limiting example, the translation menu (1130) may include an icon (1131) indicating a currently selected target language (e.g., Japanese) and an icon (1132) indicating that a translation is being performed. For example, within the UI (1110) of example (1106), while a visual effect (1169) is displayed, the icon (1132) may include an image indicating that the translation is being performed.

[0240] Referring to example (1107), the electronic device (101) may display a translated UI (1110) of the web page (or a UI (1110) changed from the UI (1110) of example (1101)) through the display (320) as the translation of the web page is completed. For example, the electronic device (101) may display translated elements of the web page within the UI (1110) of example (1107). For example, the UI (1110) of example (1107) may include a translated image (1173) including translated texts (1165) and translated text objects. For example, the translated image (1173) may be an image in which the text objects of the image (1113) are changed (or replaced) with translated text objects.

[0241] As a non-limiting example, the electronic device (101) may display a translation menu (1130) within the UI (1110) of example (1107). For example, as translation is completed, the translation menu (1130) may include an icon (1131) indicating the currently selected target language and an icon (1142) indicating that translation is complete. For example, within the UI (1110) of example (1107), as translation is completed, the display of the visual effect (1169) may stop, and the icon (1142) may include texts ('View Original') to provide a display of the web page before translation.

[0242] The electronic device (101) may display a UI (1110) representing a web page before translation upon receiving an input for an icon (1142) of the UI (1110) of the example (1107). Referring to the example (1108), the electronic device (101) may display a UI (1110) including elements of a language before translation (e.g., Korean). The UI (1110) of the example (1108) may be substantially identical to the UI (1110) of the example (1101).

[0243] As a non-limiting example, the electronic device (101) may display a modified translation menu (1130) within the UI (1110) of example (1108) upon receiving an input for the icon (1142). The translation menu (113) of example (1108) may include an icon (1131) indicating a currently selected target language and an icon (1182) indicating the display of a web page including translated elements. Although not illustrated in FIG. 11B , the electronic device (101) may re-display the UI (1110) of example (1107) upon receiving an input for the icon (1182).

[0244] Figure 12 illustrates an example of how to perform translation for a portion of a web page.

[0245] FIG. 12 illustrates examples (1201, 1202) of how an electronic device (101) performs translation of a portion of a web page.

[0246] Referring to example (1201), the electronic device (101) can display a UI (1210) (or representation) of the web page via the display (320). For example, the UI (1210) can represent at least a portion of the web page. For example, the at least portion can include texts (1211) and images (1213). For example, the texts (1211) can include texts in a first language. For example, the image (1213) can include text objects (1217) in the first language.

[0247] As a non-limiting example, the electronic device (101) may display a translation tool (1229) upon receiving (or identifying) an input for providing a translation function for a portion of the web page. Referring to example (1202), the electronic device (101) may display a translation tool (1229) shaped like a magnifying glass on the UI (1210) upon receiving the input. For example, the translation tool (1229) may be moved according to the input received with respect to the UI (1210).

[0248] Referring to example (1202), the electronic device (101) can perform a translation on an element within the UI (1210) where the translation tool (1229) is located. For example, the electronic device (101) can generate a translated image (1223) from the image (1213) by identifying the translation tool (1229) located on the image (1213). For example, the translated image (1223) can include text objects (1227) translated from the first language text objects (1217) within the image (1213) into a second language (or target language).

[0249] As a non-limiting example, the electronic device (101) may display translated elements only for portions of elements located within the translation tool (1229). For example, if a portion of text objects (1217) is located within the translation tool (1229), translated objects (1227) corresponding to said portion may be displayed within the UI (1210) of example (1202). In example (1202), texts (1211) located outside the translation tool (1229) may be maintained as texts in the first language. Since the electronic device (101) stores both the image before translation (1213) and the image after translation (1223) in the resource map by performing translation on the above web page, it can display the translated element in the translation tool (1229) instantly using the translation tool (1229) and display the element before translation again according to the movement of the translation tool (1229).

[0250] As a non-limiting example, the size of the translation tool (1229) can be adjusted, thereby allowing portions of the web page to be determined more precisely.

[0251] Figure 13 illustrates an example of how to convert a translated element of a web page into an element before translation.

[0252] FIG. 13 illustrates examples (1301, 1302) of how an electronic device (101) converts a translated element of a web page into an element before translation. In the example of FIG. 13, for convenience of explanation, the translated element is illustrated as an image, but the present disclosure is not limited thereto.

[0253] Referring to example (1301), the electronic device (101) can display a UI (1310) (or representation) of the web page through the display (320). For example, the UI (1310) can represent at least a portion of the web page. For example, the at least portion can include texts (1311) and images (1313). For example, the texts (1311) can include texts in a first language (e.g., English). For example, the image (1313) can include text objects (1317) in the first language. For example, the first language can be a target language to be translated. In other words, the UI (1310) of example (1301) can include translated elements.

[0254] As a non-limiting example, the electronic device (101) may display a translation menu (1330) within the UI (1310) of example (1301). As a non-limiting example, the translation menu (1330) may include an icon (1331) indicating a currently selected target language (e.g., English) and an icon (1332) indicating that translation is complete. For example, within the UI (1310) of example (1301), when translation is complete, the icon (1332) may include texts ('View Original') to provide a representation of the web page before translation.

[0255] When receiving an input for the icon (1332), the electronic device (101) can display elements before translation for all elements in the UI (1310). In other words, when receiving an input for the icon (1332), the electronic device (101) can display a UI that includes texts of a source language different from the first language and an image including text objects of the source language from an image (1313) including texts (1311) of the first language and text objects (1317) of the first language, and that is changed from the UI (1310).

[0256] As a non-limiting example, the electronic device (101) may receive an input (1319) for an image (1317). As a non-limiting example, the input (1319) may include a press input for the image (1317). Referring to example (1302), the electronic device (101) may display an image (1323) including pre-translation text objects (1327) from an image (1317) including translated text objects (1317) while the input (1319) is maintained. In this case, the texts (1311) within the UI (1310) may be maintained as translated texts.

[0257] In FIG. 13, an example of temporarily displaying an element before translation upon receiving an input (1319) for a UI (1310) including a translated element (or upon maintaining the input (1319)), is illustrated, but the present disclosure is not limited thereto. For example, the electronic device (101) may display a UI including a translated element upon receiving an input for the element before translation (or upon maintaining the input) for the UI including the element before translation.

[0258] Although FIGS. 5 through 13 illustrate that the electronic device (101) translates elements within the UI when a translation function for a web page is activated, the present disclosure is not limited thereto. For example, the electronic device (101) may identify elements within the UI using configuration information of the web page, and may refrain from performing translation on elements defined as a specific type (e.g., advertisement) among the elements. For example, the elements defined as the specific type may be identified based on tags in the configuration information.

[0259] Additionally, when the translation function for a web page is activated, the electronic device (101) may refrain from performing translation on elements within the UI that contain specific words. For example, the specific words may be preset for the electronic device (101) or may include proper nouns. This is because, in the case of proper nouns such as brand names, performing translation may change the original meaning.

[0260] Figure 14 illustrates an example of a flowchart for a method of generating a translated image using image and text information based on an AI model, upon determining that an image on a web page contains text.

[0261] At least some of the methods of FIG. 14 may be performed by the electronic device (101) of FIG. 3B. For example, at least some of the methods may be controlled by the processor (310) of the electronic device (101). In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0262] In operation (1410), the electronic device (101) may identify an event for translation for a web page while displaying a UI including at least a portion of a web page of an application providing a web browser.

[0263] For example, the electronic device (101) may execute the application for the web browser. For example, the application may include a software application (or a web browser application) that provides a web browser. The electronic device (101) may display a web page of the web browser on the display (320) by executing the application. For example, the electronic device (101) may display the UI of the application, which includes at least a portion of the web page, on the display (320).

[0264] For example, the electronic device (101) may identify an event for translating the web page. For example, the event may include an input to an icon representing a translation function included in a menu of the UI. However, the present disclosure is not limited thereto. For example, the event may include a designated gesture. Or, for example, the event may include an input to a pop-up message suggesting activation of the translation function. For example, the electronic device (101) may execute the translation function of the web page based on identifying the event.

[0265] For example, the electronic device (101) can identify configuration information for the web page. For example, the electronic device (101) can obtain the configuration information based on executing the translation function of the web page. For example, the electronic device (101) can obtain the configuration information for the web page using a web application programming interface (API) (e.g., selection). In the above example, the case where the configuration information is obtained based on executing the translation function (or identifying the event) is described, but the present disclosure is not limited thereto. For example, the electronic device (101) can also obtain the configuration information based on executing the application. For example, the configuration information may include code information used to configure the web page. For example, the configuration information may include HTML (hypertext markup language).

[0266] In operation (1420), the electronic device (101) can identify a first image of the web page displayed within the UI using the configuration information for the web page.

[0267] For example, the electronic device (101) may identify elements of the web page using configuration information about the web page. For example, the elements may include the first image having an image property. In the above example, the elements are described as including the first image, but the present disclosure is not limited thereto. For example, the elements may further include elements having a text property (e.g., texts) or elements having an image property (e.g., other images).

[0268] In operation (1430), the electronic device (101) may determine whether the first image includes text based on the AI ​​model (340) using the first image. For example, the electronic device (101) may determine whether the first image includes text based on the AI ​​model (340). For example, the electronic device (101) may output a first value based on the AI ​​model (340) upon determining that the first image includes text. For example, the electronic device (101) may output a second value based on the AI ​​model (340) upon determining that the first image does not include text. In the following, for convenience of explanation, it is assumed that the first image includes a first text, and the first text is in a source language different from a target language.

[0269] In operation (1440), the electronic device (101) may obtain text information of the first text based on the AI ​​model (340) upon determining that the first image includes the first text.

[0270] For example, the text information may include an OCR result of the first text in the first image (or contents or characters of the first text), location information (or coordinates) of the first text in the first image, or size information of a virtual block where the first text is located. However, the present disclosure is not limited thereto. For example, the text information may further include rotation information indicating a degree to which the first text in the first image is rotated with respect to a specified reference.

[0271] In operation (1450), the electronic device (101) may generate a second image including a second text translated from the first text into the target language based on an AI model (340) using the text information of the first text and the first image.

[0272] For example, the electronic device (101) may perform a translation on the first image and generate the translated second image. For example, performing the translation on the first image may include translating the first text within the image into text in a target language. For example, if the language (or source language) of the first text within the first image is different from the target language, the translation of the first text may be performed. As a non-limiting example, if the language (or source language) of the third text within the first image is the same as the target language, the translation of the third text may not be performed.

[0273] In one example, the electronic device (101) can perform translations of first texts in a first source language among the texts in the first image substantially simultaneously based on the AI ​​model (340). For example, the first texts may include the first text. After performing translations of the first texts, the electronic device (101) can perform translations of second texts in a second source language among the texts in the first image substantially simultaneously based on the AI ​​model (340). In other words, the electronic device (101) can perform translations based on the AI ​​model (340) for texts having the same source language among the texts in the first image.

[0274] For example, the electronic device (101) may generate the second text translated into the target language, and then generate the translated second image including the second text. In one example, the electronic device (101) may adjust the properties of the second text to correspond to the properties of the first image based on the AI ​​model (340). For example, the electronic device (101) may identify the properties of the first image based on the AI ​​model (340). For example, the properties of the first image may include the location of the first image, the locations and colors of elements (e.g., images) included in the first image, the representative color of the first image, or the size of the first image. For example, the electronic device (101) may adjust the properties of the translated second text according to the representative color of the first image. For example, the properties of the text may include at least one of the color, thickness, size, font, and background color of the text.

[0275] For example, the electronic device (101) may store the second image in a resource map (e.g., the resource map (950) of FIG. 9). For example, the resource map may be stored in a storage area used by the application. For example, the electronic device (101) may store the resource map for one or more elements of the web page in the storage area as the application for the web browser is executed. For example, the resource map may include image information of an element having an image property displayed in the web page (or information on pixels constituting an image). For example, the resource map may include the first image including the first text before being translated. For example, the electronic device (101) may store the second image in the resource map. For example, the electronic device (101) may store both the first image and the second image in the resource map.

[0276] For example, the electronic device (101) can display the UI including the second image. For example, the electronic device (101) can display the UI including the second image changed from the first image on the display (320) by identifying the second image stored in the resource map. In other words, the electronic device (101) can generate and store the second image based on identifying the event, and then display the UI including the second image replaced from the first image by identifying the second image instead of the first image in the resource map.

[0277] Referring to FIGS. 1 to 14, the present disclosure can determine an image including text among elements having image properties of a web page based on an AI model (340) stored in an electronic device (101), and perform a translation on the image including text. For example, the present disclosure can perform a translation on an image including text based on the AI ​​model (340), and provide a naturally translated web page by adjusting the properties of the translated text to be similar to the properties of the image. In addition, the present disclosure can define the order (or method) of the translation on elements of the web page performed based on the AI ​​model (340). Accordingly, the present disclosure can perform the translation performed based on the AI ​​model (340) more quickly. In addition, the present disclosure can store the translated image generated by performing the translation based on the AI ​​model (340) and the translated text included in the translated image in a separate storage space (or resource map) related to the web page. For example, the present disclosure can provide a user with a translated web page more quickly by using translated images and translated text stored in the separate storage space when the web page is changed or the translation function of the web page is activated and deactivated.

[0278] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs from the description below.

[0279] As described above, the electronic device (101) may include a memory (330) that stores instructions and includes one or more storage media. The electronic device (101) may include at least one processor (310) that includes a processing circuit. The electronic device (101) may include a display (320). The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify an event for translation of a web page while displaying a user interface (UI) including at least a portion of a web page of an application providing a web browser on the display (320). The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify a first image of the web page displayed within the UI by using configuration information about the web page upon identifying the event. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to determine, based on an artificial intelligence (AI) model stored in the electronic device (101), whether the first image includes text using the first image. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain text information of the first text based on the AI ​​model upon determining that the first image includes the first text.The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate a second image including a second text translated into a target language from the first text based on the AI ​​model using the text information of the first text and the first image. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, on the display (320), the UI of the application including the second image modified from the first image.

[0280] According to one embodiment, the properties of the second text of the second image may be adjusted based on the representative colors of the pixels of the first image, based on the AI ​​model. The properties may include at least one of the color, size, or font of the text.

[0281] In one embodiment, the event for the translation may include an input to an icon representing a translation function included in a menu of the UI.

[0282] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify a first region corresponding to the at least a portion of the web page displayed on the display (320) through the UI, and a second region extending from the first region and not displayed on the display (320), upon identifying the event. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify the first image of the first region using the configuration information for the web page.

[0283] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify, within the first region, third text distinct from the first image using the configuration information for the web page. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate, based on the AI ​​model, fourth text translated from the third text into the target language using the third text before determining whether the first image includes text. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, on the display (320), the UI of the application including the fourth text modified from the third text and the second image.

[0284] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify a plurality of images including the first image of the first area using the configuration information for the web page. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify a first image having a size exceeding a reference size and a third image having a size less than or equal to the reference size among the plurality of images. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to determine, based on the AI ​​model, using the third image, whether the third image includes text. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain text information of the fifth text based on the AI ​​model using the third image, upon determining that the third image includes a fifth text. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to, after generating the second text, generate a fourth image including a sixth text translated into the target language from the fifth text based on the AI ​​model using the text information of the fifth text and the third image.The above instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, on the display (320), the UI of the application, which includes the second image changed from the first image and the fourth image changed from the third image.

[0285] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify a fifth image of the second area using the configuration information for the web page. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to, after generating the second image and the fourth image, determine, based on the AI ​​model, whether the fifth image of the second area includes text using the fifth image of the second area. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain text information of the seventh text based on the AI ​​model upon determining that the fifth image of the second area includes a seventh text. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate a sixth image including an eighth text translated from the seventh text into the target language based on the AI ​​model using the text information of the seventh text and the fifth image. The sixth image of the second area may be displayed on the display (320) according to a user input received with respect to the UI.

[0286] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify a plurality of images including the first image of the first area using the configuration information for the web page. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to perform a translation on a first set of images among the plurality of images, the first set corresponding to the reference number and including the first image, based on the AI ​​model, upon determining that the number of the plurality of images exceeds a reference number. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to perform a translation on a second set of images among the plurality of images, the second set of images being different from the first set of images, based on the AI ​​model, after performing the translation on the first set of images.

[0287] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to refrain from obtaining the text information of the first text based on the AI ​​model using the first image, upon determining that the first image does not include text.

[0288] According to one embodiment, the first image may include the first text, the ninth text, and the tenth text. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain text information of the ninth text and text information of the tenth text based on the AI ​​model, upon determining that the first image includes the ninth text and the tenth text. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to use the text information of the first text, the text information of the ninth text, and the text information of the tenth text to identify the first text in a first language different from the target language and the ninth text in the first language, and the tenth text in a second language different from each of the target language and the first language. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain, based on the AI ​​model, the second text translated into the target language from the first text and the eleventh text translated into the target language from the tenth text, using the text information of the first text and the text information of the tenth text. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain, based on the AI ​​model, the twelfth text translated into the target language from the ninth text, using the text information of the ninth text, after obtaining the second text and the eleventh text.The above instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate the second image including the second text, the eleventh text, and the twelfth text.

[0289] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to execute the application for the web browser. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to store a resource map for one or more elements including the first image of the web page in a storage area utilized by the application. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate the second image and then store the second image in the resource map of the storage area.

[0290] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify another event within the UI for re-displaying the first image while displaying the UI of the application including the second image. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, on the display (320), the UI including the first image changed from the second image by identifying the first image within the resource map upon identifying the other event. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, on the display (320), the UI including the second image changed from the first image by identifying the second image in the resource map, upon re-identifying the event for the translation of the web page.

[0291] In one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to receive an input regarding the UI while displaying the UI including the second image on the display (320) upon identifying the other event. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to, upon receiving the input, display the UI including at least a portion of another web page modified from the web page on the display (320). The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to delete the resource map including the second image from the storage area upon displaying the UI including the at least a portion of the other web page on the display (320).

[0292] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to store the text information of the first text of the first image within the electronic device (101) when deleting the resource map within the storage area. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to receive another input with respect to the UI while displaying the UI including the at least a portion of the other web page. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, on the display (320), the UI including the first image of the web page changed from the other web page upon receiving the another input. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to: display the UI including the first image of the web page upon receiving the other input; then, upon identifying an event for translation of the web page, skip determining whether the first image includes text based on the AI ​​model using the first image; and re-generate the second image based on the AI ​​model using the text information of the first text stored in the electronic device (101).

[0293] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to, upon receiving the input, display a notification object on the display (320) that inquires whether to store the second image within the resource map in another storage area of ​​the electronic device (101) distinct from the storage area. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to, upon receiving another input for the notification object, store the second image within the resource map in the other storage area of ​​the electronic device (101), and delete the resource map including the second image from the storage area.

[0294] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate additional information together with the second image based on the AI ​​model using the text information of the first text and the first image. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, on the display (320), the UI including a visual object representing the second image and the additional information. The additional information may include a summary of the first text.

[0295] According to one embodiment, the text information of the first text may include content of the first text, location information of the first text within the first image, and size information of a virtual block in which the first text is located.

[0296] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain configuration information for the web page upon identifying the event. The configuration information may include code information for elements within the web page. The elements may include text and images within the web page.

[0297] As described above, the method performed by the electronic device (101) may include an operation of identifying an event for translating a web page while displaying a user interface (UI) including at least a portion of a web page of an application providing a web browser. The method may include an operation of identifying a first image of the web page displayed in the UI using configuration information about the web page upon identifying the event. The method may include an operation of determining, using the first image, whether the first image includes text based on an artificial intelligence (AI) model stored in the electronic device (101). The method may include an operation of obtaining text information of the first text based on the AI ​​model upon determining that the first image includes the first text. The method may include an operation of generating, using the text information of the first text and the first image, a second image including second text translated into a target language from the first text based on the AI ​​model. The method may include an action of displaying the UI of the application including the second image changed from the first image.

[0298] The non-transitory computer-readable storage medium as described above may store one or more programs including instructions that, when individually or collectively executed by at least one processor (310) of an electronic device (101) including a display (320), cause the electronic device (101) to identify an event for translation of a web page while displaying a user interface (UI) including at least a portion of a web page of an application providing a web browser on the display (320). The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to identify a first image of the web page displayed within the UI using configuration information about the web page upon identifying the event. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to determine, based on an artificial intelligence (AI) model stored in the electronic device (101), whether the first image includes text using the first image. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to obtain text information of the first text based on the AI ​​model upon determining that the first image includes the first text.The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to generate a second image including a second text translated into a target language from the first text based on the AI ​​model using the text information of the first text and the first image. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to display, on the display (320), the UI of the application including the second image changed from the first image.

[0299] As described above, the electronic device (101) may include a memory (330) that includes one or more storage media and stores instructions. The electronic device (101) may include at least one processor (310) that includes a processing circuit. The electronic device (101) may include a display (320). The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to display a first representation of a web page via the display (320). The first representation may include a first image. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify an input for performing a translation from a first language to a second language with respect to the web page. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to perform optical character recognition (OCR) of the first image to obtain first text objects in the first language from the first image based on the input. Information about the arrangement of the first text objects within the first image may be further identified by performing the OCR of the first image. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain second text objects in the second language by performing the translation from the first language to the second language on the obtained first text objects.The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify arrangement information of the second text objects in a second image to be generated based on the arrangement information of the first text objects. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate a second image including the second text objects by inpainting the second text objects into the first image using the arrangement information of the second text objects. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, through the display (320), a second representation of the web page including the second image.

[0300] According to one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify a representative color of the first image. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate a second image including the second text objects based on applying the representative color of the first image to the second text objects of the second image.

[0301] In one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to provide the first image to a trained model running on the electronic device (101) based on the input. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to determine, using the trained model, whether the first image is an image having a text object. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain the first text objects and the arrangement information of the first text objects within the first image by performing the OCR of the first image using the trained model based on determining that the first image is an image having a text object. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain the second text objects by performing the translation from the first language to the second language using the trained model based on providing the obtained first text objects to the trained model. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate the second image including the second text objects using the trained model based on providing the first image and the second text objects to the trained model.

[0302] According to one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify, based on the input, a first region of the web page that is displayed through the display (320) and corresponds to the first representation of the web page, and a second region that is not displayed through the display (320) and extends from the first region. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify, using a markup language file of the web page, the first image of the first region.

[0303] In one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify, within the first region, first texts in the first language distinct from the first image, using the markup language file of the web page. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to provide the first texts in the first language to the trained model. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate second texts in the second language by performing, using the trained model, the translation from the first language to the second language with respect to the web page before determining whether the first image is an image having a text object. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, via the display (320), the second representation of the web page including the second texts and the second image.

[0304] According to one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify a plurality of images including the first image of the first area using the markup language file of the web page. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify the first image among the plurality of images having a size exceeding a reference size. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to perform the OCR of the first image based on the input.

[0305] In one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify a third image of the second area using the markup language file of the web page. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to, after generating the second image, determine, using the trained model, whether the third image of the second area is an image having a text object. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain third text objects of the first language within the third image by performing the OCR of the third image using the trained model based on determining that the third image of the second area is an image having a text object. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain fourth text objects in the second language by performing the translation from the first language to the second language using the trained model based on providing the obtained third text objects to the trained model. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to provide the third image and the fourth text objects to the trained model.The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate a fourth image including the fourth text objects based on removing the third text objects within the third image and inpainting the fourth text objects within the third image. The fourth image of the second region may be displayed via the display (320) in response to input received regarding the second representation of the web page displayed via the display (320).

[0306] In one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify, using the markup language file of the web page, a plurality of images including the first image of the first area. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to perform, using the trained model, a translation into the second language for a first image set among the plurality of images that corresponds to the reference number and includes the first image, based on determining that the number of the plurality of images exceeds a reference number. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to perform the translation into the second language for the first set of images, and then, using the trained model, perform the translation into the second language for a second set of images, different from the first set of images, among the plurality of images.

[0307] According to one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to store at least one of the first text objects or the second text objects in a clipboard running on the electronic device (101) based on obtaining the second text objects.

[0308] According to one embodiment, the first image may include the first text objects, the fifth text objects, and the sixth text objects. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain the fifth text objects and the sixth text objects by performing the OCR of the first image using the trained model based on determining that the first image is an image having a text object. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify the first text objects in the first language different from the second language and the fifth text objects in the first language, and the sixth text objects in a third language different from each of the second language and the first language. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain the second text objects and the seventh text objects by performing the translation from the first language to the second language using the trained model based on providing the obtained first text objects and the obtained fifth text objects to the trained model.The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain eighth text objects by performing a translation from the third language to the second language using the trained model based on obtaining the second text objects and the seventh text objects and then providing the obtained sixth text objects to the trained model. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate the second image including the second text objects, the seventh text objects, and the eighth text objects based on removing the fifth text objects within the first image and inpainting the seventh text objects within the first image, and removing the sixth text objects within the first image and inpainting the eighth text objects within the first image.

[0309] According to one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to execute an application for a web browser providing the web page. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to store a resource map for one or more elements including the first image of the web page in a storage area utilized by the application. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate the second image and then store the second image in the resource map of the storage area.

[0310] In one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify another input for re-displaying the representation of the web page including the first image while displaying the second representation of the web page including the second image. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, through the display (320), the first representation of the web page including the first image by identifying, based on the another input, the first image stored in the resource map. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, through the display (320), the second representation of the web page including the second image by identifying the second image within the resource map based on the input for the translation of the web page.

[0311] In one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to receive another input regarding the second representation of the web page while displaying the UI including the second image on the display (320). The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, through the display (320), a third representation of another web page modified from the web page based on the another input. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to remove the resource map including the first image and the second image from the storage area based on displaying the third representation of the other web page via the display (320).

[0312] According to one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to store the first text objects of the first image within the electronic device (101) when removing the resource map from the storage area. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to receive input regarding the third representation of the other web page while displaying the third representation of the other web page. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, through the display (320), the first representation of the web page including the first image based on the input regarding the third representation of the other web page. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to skip determining, using the trained model, whether the first image is an image having text objects by providing the first image to a trained model executed on the electronic device (101) based on identifying an input for performing a translation of the web page.The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate the second image using the trained model based on identifying an input for performing a translation of the web page, and based on providing the first text objects stored in the electronic device (101) back to the trained model.

[0313] In one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate, using the trained model, additional information together with the second image based on providing the first text objects and the first image to the trained model. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, via the display (320), the second representation of the web page, including the second image and a visual object representing the additional information. The additional information may include a summary of the first text objects.

[0314] According to one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain, by using the trained model, text information including the first text objects, the arrangement information of the first text objects within the first image, and the size information of a virtual block in which the first text objects are located, by performing the OCR of the first image using the trained model.

[0315] In one embodiment, the second image may be generated by removing the first text objects within the first image, inpainting the second text objects into a portion of an area of ​​the first image containing the first text objects, and inpainting another portion of the area of ​​the first image with a color identified from the first image.

[0316] According to one embodiment, the instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify an input with respect to the second image while displaying the second representation of the web page including the second image. The instructions, when individually and / or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, through the display (320), the second representation of the web page including the first image modified from the second image while the identified input with respect to the second image is maintained.

[0317] A method performed by an electronic device (101) as described above may include an operation of displaying a first representation of a web page. The first representation may include a first image. The method may include an operation of identifying an input for performing a translation from a first language to a second language with respect to the web page. The method may include an operation of performing optical character recognition (OCR) on the first image to obtain first text objects in the first language from the first image based on the input. Placement information of the first text objects within the first image may be further identified by performing the OCR on the first image. The method may include an operation of performing the translation from the first language to the second language with respect to the obtained first text objects, thereby obtaining second text objects in the second language. The method may include an operation of identifying, based on the placement information of the first text objects, placement information of the second text objects within a second image to be generated. The method may include an action of generating the second image including the second text objects by inpainting the second text objects onto the first image using the arrangement information of the second text objects. The method may include an action of displaying a second representation of the web page including the second image.

[0318] The non-transitory computer-readable storage medium as described above may store one or more programs including instructions that, when individually or collectively executed by at least one processor (310) of an electronic device (101) including a display (320), cause the electronic device (101) to display a first representation of a web page via the display (320). The first representation may include a first image. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to identify an input for performing a translation from a first language to a second language with respect to the web page. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to perform optical character recognition (OCR) of the first image to obtain first text objects in the first language from the first image based on the input. Information about the arrangement of the first text objects within the first image may be further identified by performing the OCR of the first image. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to perform the translation from the first language to the second language on the obtained first text objects, thereby obtaining second text objects in the second language.The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to identify arrangement information of the second text objects in a second image to be generated based on the arrangement information of the first text objects. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to generate the second image including the second text objects by inpainting the second text objects into the first image using the arrangement information of the second text objects. The non-transitory computer-readable storage medium may store one or more programs comprising instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to display, through the display (320), a second representation of the web page including the second image.

[0319]

[0320] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0321] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0322] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0323] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0324] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0325] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device (101), A memory (330) comprising one or more storage media and storing instructions; At least one processor (310) comprising a processing circuit; and Includes a display (320), The above instructions, when individually and / or collectively executed by the at least one processor (310): Displaying a first representation of a web page through the display (320), wherein the first representation includes a first image; Identifying input for performing translation from a first language to a second language with respect to said web page; Based on the input, performing optical character recognition (OCR) of the first image to obtain first text objects of the first language from the first image, and arrangement information of the first text objects within the first image is further identified by performing the OCR of the first image; Obtaining second text objects in the second language by performing the translation from the first language to the second language with respect to the first text objects obtained above; Based on the arrangement information of the first text objects, identify the arrangement information of the second text objects in the second image to be generated; Generating the second image including the second text objects by inpainting the second text objects into the first image using the arrangement information of the second text objects; and To display a second representation of the web page including the second image through the display (320), The above electronic device (101), causing, Electronic device (101).

2. In claim 1, The above instructions, when individually and / or collectively executed by the at least one processor (310): Identifying the representative color of the first image; and To generate the second image including the second text objects based on applying the representative color of the first image to the second text objects of the second image, The above electronic device (101), causing, Electronic device (101).

3. In claim 1 The above instructions, when individually and / or collectively executed by the at least one processor (310): Based on the above input, providing the first image to the trained model running on the electronic device (101); Using the trained model, determine whether the first image is an image having a text object; Based on determining that the first image is an image having a text object, performing the OCR of the first image using the trained model to obtain the first text objects within the first image and the arrangement information of the first text objects; Obtaining the second text objects by performing the translation from the first language to the second language using the trained model based on providing the acquired first text objects to the trained model; and Based on providing the first image and the second text objects to the trained model, generate the second image including the second text objects using the trained model. The above electronic device (101), causing, Electronic device (101).

4. In claim 3, The above instructions, when individually and / or collectively executed by the at least one processor (310): Based on the input, identifying a first area displayed through the display (320) among the web pages and corresponding to the first expression of the web page, and a second area extending from the first area and not displayed through the display (320); and To identify the first image of the first area using a markup language file of the web page, The above electronic device (101), causing, Electronic device (101).

5. In claim 4, The above instructions, when individually and / or collectively executed by the at least one processor (310): Using the markup language file of the web page, identifying first texts of the first language that are distinct from the first image within the first area; providing the first texts of the first language to the trained model; and Before determining whether the first image is an image having a text object, generating second texts in the second language by performing the translation from the first language to the second language with respect to the web page using the trained model; and To display the second representation of the web page including the second texts and the second image through the display (320), The above electronic device (101) causes Electronic device (101).

6. In claim 4, The above instructions, when individually and / or collectively executed by the at least one processor (310): Using the markup language file of the web page, identifying a plurality of images including the first image in the first area; Identifying the first image among the plurality of images having a size exceeding a reference size; and Based on the above input, to perform the OCR of the first image, The above electronic device (101), causing, Electronic device (101).

7. In claim 4, The above instructions, when individually and / or collectively executed by the at least one processor (310): Using the markup language file of the above web page, identify the third image of the second area; After generating the second image, using the trained model, it is determined whether the third image in the second area is an image having a text object; Based on determining that the third image of the second area is an image having a text object, performing the OCR of the third image using the trained model to obtain third text objects of the first language within the third image; and By providing the acquired third text objects to the trained model, the translation from the first language to the second language is performed using the trained model, thereby acquiring fourth text objects of the second language; Providing the third image and the fourth text objects to the trained model; To generate a fourth image including the fourth text objects based on removing the third text objects within the third image and inpainting the fourth text objects within the third image, The above electronic device (101) causes, The fourth image of the second area is to be displayed through the display (320) according to an input received regarding the second representation of the web page displayed through the display (320). Electronic device (101).

8. In claim 4, The above instructions, when individually and / or collectively executed by the at least one processor (310): Using the markup language file of the web page, identifying a plurality of images including the first image in the first area; Based on determining that the number of the plurality of images exceeds a reference number, using the trained model, performing a translation into the second language for a first image set that corresponds to the reference number among the plurality of images and includes the first image; and After performing the translation into the second language for the first image set, using the trained model, perform the translation into the second language for a second image set different from the first image set among the plurality of images. The above electronic device (101), causing, Electronic device (101).

9. In claim 3, The above instructions, when individually and / or collectively executed by the at least one processor (310): Based on obtaining the second text objects, storing at least one of the first text objects or the second text objects in a clipboard running on the electronic device (101). The above electronic device (101), causing, Electronic device (101).

10. In claim 1, The first image includes the first text objects, the fifth text objects, and the sixth text objects, The above instructions, when individually and / or collectively executed by the at least one processor (310), cause the electronic device (101) to: Based on determining that the first image is an image having a text object, performing the OCR of the first image using the trained model to obtain the fifth text objects and the sixth text objects; Identifying the first text objects of the first language different from the second language, the fifth text objects of the first language, and the sixth text objects of the third language different from each of the second language and the first language; By providing the acquired first text objects and the acquired fifth text objects to the trained model, the translation from the first language to the second language is performed using the trained model, thereby acquiring the second text objects and the seventh text objects; After obtaining the second text objects and the seventh text objects, by performing a translation from the third language to the second language using the trained model based on providing the obtained sixth text objects to the trained model, thereby obtaining the eighth text objects; and Generate the second image including the second text objects, the seventh text objects, and the eighth text objects based on removing the fifth text objects in the first image and inpainting the seventh text objects in the first image, and removing the sixth text objects in the first image and inpainting the eighth text objects in the first image. The above electronic device (101), causing, Electronic device (101).

11. In claim 1, The above instructions, when individually and / or collectively executed by the at least one processor (310): Run an application for a web browser that provides the above web page; storing a resource map for one or more elements including the first image of the web page in a storage area used by the application; and After generating the second image, store the second image within the resource map of the storage area. The above electronic device (101), causing, Electronic device (101).

12. In claim 11, The above instructions, when individually and / or collectively executed by the at least one processor (310): While displaying said second representation of said web page including said second image, identifying another input for displaying said representation of said web page including said first image again; Based on the other input, by identifying the first image stored in the resource map, displaying the first representation of the web page including the first image through the display (320); and Based on the input for the translation of the web page, by identifying the second image in the resource map, displaying the second representation of the web page including the second image through the display (320). The above electronic device (101), causing, Electronic device (101).

13. In claim 11, The above instructions, when individually and / or collectively executed by the at least one processor (310): While displaying the UI including the second image on the display (320), receiving another input regarding the second representation of the web page; Based on the other input, displaying a third representation of another web page changed from the web page through the display (320); and Based on displaying the third representation of the other web page through the display (320), to remove the resource map including the first image and the second image from the storage area, The above electronic device (101), causing, Electronic device (101).

14. In a method performed by an electronic device (101), the method: An act of displaying a first representation of a web page, said first representation including a first image; An action to identify an input for performing a translation from a first language to a second language with respect to the above web page; An operation of performing optical character recognition (OCR) of the first image to obtain first text objects of the first language from the first image based on the input, wherein arrangement information of the first text objects within the first image is further identified by performing the OCR of the first image; An operation of obtaining second text objects of the second language by performing the translation from the first language to the second language with respect to the obtained first text objects; An operation of identifying arrangement information of the second text objects in a second image to be generated based on the arrangement information of the first text objects; An operation of generating the second image including the second text objects by inpainting the second text objects into the first image using the arrangement information of the second text objects; and comprising an action of displaying a second representation of the web page including the second image; method.

15. In a non-transitory computer-readable storage medium, when individually or collectively executed by at least one processor (310) of an electronic device (101) including a display (320), the electronic device (101): Displaying a first representation of a web page through the display (320), wherein the first representation includes a first image; Identifying input for performing translation from a first language to a second language with respect to said web page; Based on the input, performing optical character recognition (OCR) of the first image to obtain first text objects of the first language from the first image, and arrangement information of the first text objects within the first image is further identified by performing the OCR of the first image; Obtaining second text objects in the second language by performing the translation from the first language to the second language with respect to the first text objects obtained above; Based on the arrangement information of the first text objects, identify the arrangement information of the second text objects in the second image to be generated; Generating the second image including the second text objects by inpainting the second text objects into the first image using the arrangement information of the second text objects; and storing one or more programs including instructions that cause a second representation of the web page including the second image to be displayed through the display (320); Non-transitory computer-readable storage medium.

Citation Information

Patent Citations

  • Translation display device

    JP2012118959A

  • Image processing apparatus and image forming apparatus

    JP2019117988A

  • Smart translation system

    JP2023155158A

  • Method, apparatus and computer-readable recording medium for reading text on image contained in web page and providing translation service on same text

    KR100953627B1

  • Transparent frame and manufacturing method therefor

    KR102722686B1