Electronic device for supporting image analysis, operation method thereof, and storage medium
Patent Information
- Application Number
- PCT/KR2026/002672
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-30
- Filing Date
- 2026-02-12
- Publication Date
- 2026-10-01
Smart Images

Figure KR2026002672_01102026_PF_FP_ABST
Abstract
Description
Electronic device supporting image analysis, method of operation thereof, and storage medium
[0001] The present disclosure relates to an electronic device that supports image analysis, a method of operation thereof, and a storage medium.
[0002] Optical Character Recognition (OCR) technology can be a technology that analyzes images to detect character regions and converts them into character data. OCR technology can be used in various application fields, including document digitization, automated data entry, and augmented reality (AR). Scene Text Recognition (STR), based on OCR technology, is a technology that detects and reads text from captured images and / or videos, and can be utilized for road sign recognition, automated document processing, and / or augmented reality-based information provision.
[0003] With the advancement of artificial intelligence (AI) technology, various models that learn using massive amounts of data are emerging. For example, Large Language Models (LLMs) are models specialized for natural language processing that can provide natural language understanding and / or generation capabilities by learning from large volumes of text data. For instance, Retrieval-Augmented Generation (RAG) techniques can refine the responses of existing language models by retrieving relevant information from external databases. In the field of image processing, for instance, Large Vision Models (LVMs) have been developed to perform various functions such as image recognition, object detection, and / or context awareness. For instance, Large World Models (LWMs) can understand and predict complex worlds by learning based on simulations of physical environments.
[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0005] The electronic device may include at least one processor and a memory for storing instructions.
[0006] When the above instructions are executed individually or collectively by the at least one processor, they may cause the electronic device to provide a first image.
[0007] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify a first object on which a first text group included in the first image is written and a second object on which a second text group is written, based on the confirmation of a request for text analysis for the first image.
[0008] When the above instructions are executed individually or collectively by the at least one processor, they may cause the electronic device to verify the physical arrangement relationship in the real world between the first object and the second object.
[0009] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide an analysis result associated with a third text group including the first text group and the second text group, based on the determination that the first object and the second object together constitute an object in which a text group is written, based on the physical arrangement relationship.
[0010] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide an analysis result associated with either the first text group or the second text group based on the determination that the first object and the second object correspond to objects on which different text groups are written, based on the physical arrangement relationship.
[0011] The method of operation of an electronic device may include an operation that provides a first image.
[0012] The method of operation of the electronic device may include, based on the confirmation of a request for text analysis regarding the first image, an operation of confirming a first object on which a first text group included in the first image is written and a second object on which a second text group is written.
[0013] The method of operation of the electronic device may include an operation to verify the physical arrangement relationship in the real world between the first object and the second object.
[0014] The method of operation of the electronic device may include providing an analysis result associated with a third text group including the first text group and the second text group based on the determination that the first object and the second object together constitute an object in which a single text group is written based on the physical arrangement relationship, or providing an analysis result associated with either the first text group or the second text group based on the determination that the first object and the second object each correspond to objects in which different text groups are written based on the physical arrangement relationship.
[0015] A storage medium for storing computer-readable instructions may be provided.
[0016] When the above instructions are executed individually or collectively by at least one processor of the electronic device, the electronic device may cause the electronic device to provide a first image.
[0017] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify a first object on which a first text group included in the first image is written and a second object on which a second text group is written, based on the confirmation of a request for text analysis for the first image.
[0018] When the above instructions are executed individually or collectively by the at least one processor, they may cause the electronic device to verify the physical arrangement relationship in the real world between the first object and the second object.
[0019] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide an analysis result associated with a third text group including the first text group and the second text group, based on the determination that the first object and the second object together constitute an object in which a text group is written, based on the physical arrangement relationship.
[0020] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide an analysis result associated with either the first text group or the second text group based on the determination that the first object and the second object correspond to objects on which different text groups are written, based on the physical arrangement relationship.
[0021] The electronic device may include at least one processor and a memory for storing instructions.
[0022] When the above instructions are executed individually or collectively by the at least one processor, they may cause the electronic device to provide a first image.
[0023] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify a first object on which a first text group is written and a second object on which a second text group is written, based on the confirmation of a request for text analysis for the first image.
[0024] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify at least a portion of the first text group or the second text group as text to be recognized, based on the physical arrangement relationship between the first object and the second object.
[0025] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause to provide an analysis result associated with the recognition target text and a second image including a third object generated based on adjustments to at least some of the first object and the second object.
[0026] The method of operation of an electronic device may include an operation of providing a first image.
[0027] The method of operation of the electronic device may include an operation of verifying a first object on which a first text group is written and a second object on which a second text group is written, based on the verification of a request for text analysis for the first image.
[0028] The method of operation of the electronic device may include an operation of identifying at least a portion of the first text group or the second text group as a text to be recognized, based on the physical arrangement relationship between the first object and the second object.
[0029] The method of operation of the electronic device may include providing an analysis result associated with a second image comprising a third object generated based on adjustment of at least some of the first object and the second object, and the text to be recognized.
[0030] A storage medium for storing computer-readable instructions may be provided.
[0031] When the above instructions are executed individually or collectively by at least one processor of the electronic device, the electronic device may cause the electronic device to provide a first image.
[0032] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify a first object on which a first text group is written and a second object on which a second text group is written, based on the confirmation of a request for text analysis for the first image.
[0033] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify at least a portion of the first text group or the second text group as text to be recognized, based on the physical arrangement relationship between the first object and the second object.
[0034] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause to provide an analysis result associated with the recognition target text and a second image including a third object generated based on adjustments to at least some of the first object and the second object.
[0035] The electronic device may include at least one processor and a memory for storing instructions.
[0036] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause to provide a first image comprising a first object on which a first text group is written and a second object on which a second text group is written.
[0037] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause to provide a second image comprising a third object generated based on adjustments to at least some of the first object and the second object, and a third text group written on the third object, and an analysis result associated with the third text group, based on the confirmation of a request for text analysis for the first image.
[0038] Here, the third text group may be generated based on at least some of the first text group or the second text group.
[0039] A method of operation of an electronic device may include an operation of providing a first image comprising a first object on which a first text group is written and a second object on which a second text group is written.
[0040] The method of operation of the electronic device may include, based on the confirmation of a request for text analysis regarding the first image, providing a second image comprising a third object generated based on adjustment of at least some of the first object and the second object, and a third text group written on the third object, and an analysis result associated with the third text group.
[0041] Here, the third text group may be generated based on at least some of the first text group or the second text group.
[0042] A storage medium for storing computer-readable instructions may be provided.
[0043] When the above instructions are executed individually or collectively by at least one processor of the electronic device, the electronic device may cause to provide a first image comprising a first object on which a first text group is written and a second object on which a second text group is written.
[0044] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause to provide a second image comprising a third object generated based on adjustments to at least some of the first object and the second object, and a third text group written on the third object, and an analysis result associated with the third text group, based on the confirmation of a request for text analysis for the first image.
[0045] Here, the third text group may be generated based on at least some of the first text group or the second text group.
[0046] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0047] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments.
[0048] FIGS. 2A and 2B are drawings for explaining the operation method of an electronic device according to a comparative example for comparison with embodiments.
[0049] FIG. 3 is a diagram illustrating a method of operation of an electronic device according to one embodiment.
[0050] Figures 4a, 4b, 4c, and 4d are drawings for explaining the provision of analysis results associated with text.
[0051] FIG. 4e is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0052] FIG. 5 is a diagram illustrating a method of operation of an electronic device according to one embodiment.
[0053] FIGS. 6a, FIGS. 6b, FIGS. 6c and FIGS. 6d are examples of screens displayed by an electronic device according to various embodiments.
[0054] FIG. 7 is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0055] FIG. 8 is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0056] FIG. 9 is a diagram illustrating the classification results according to one embodiment.
[0057] FIG. 10 is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0058] FIG. 11a is a diagram illustrating the determination of a text to be recognized based on a user selection according to one embodiment.
[0059] FIG. 11b is a drawing for illustrating an image provided by an electronic device according to one embodiment.
[0060] FIG. 12a is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0061] FIG. 12b is a drawing for illustrating an image provided by an electronic device according to one embodiment.
[0062] FIG. 13 is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0063] FIGS. 14a, 14b, and 14c are drawings for illustrating images provided by an electronic device according to one embodiment.
[0064] FIG. 15 is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0065] FIGS. 16a, FIGS. 16b, and FIGS. 16c are drawings for illustrating text analysis based on objects not associated with text.
[0066] FIG. 17 is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0067] FIGS. 18a and FIGS. 18b are drawings for illustrating text analysis based on objects not associated with text.
[0068] FIG. 19 is a drawing for explaining the operation of an electronic device according to one embodiment.
[0069] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.
[0070] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments.
[0071] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0072] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0073] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0074] The number of processors (120) may be one or more. For example, the processor (120) may have the structure of a multi-core processor such as a dual core, a quad core, or a hexa core.
[0075] The processor (120) can control the operations of the electronic device (101) by executing instructions stored in memory (130). For example, the processor (120) may correspond to a plurality of processors that divide and collectively perform a plurality of operations among the processors.
[0076] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0077] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0078] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0079] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0080] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0081] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0082] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0083] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0084] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0085] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0086] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0087] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0088] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0089] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0090] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0091] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0092] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0093] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0094] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0095] FIGS. 2a and 2b are drawings illustrating a method of operation of an electronic device according to a comparative example for comparison with the embodiments. Those skilled in the art will understand that some or all of the operations according to the comparative example may be performed by the embodiments.
[0096] Referring to FIG. 2a, in the real world, a first document (201) may be placed on a second document (202), and the second document (202) may be placed on a floor (203). A first image (210) may be obtained by taking a picture of the documents (201, 202) placed on the floor (203) in the real world as in FIG. 2b. Referring to FIG. 2b, the first image (210) may include a first object (211) corresponding to the first document (201) in the real world, a second object (212) corresponding to the second document (202), and a third object (213) corresponding to the floor (203). The first object (211) may include a first text group, and the second object (212) may include a second text group. Those skilled in the art will understand that in the present disclosure, a first text group may be written on a first object (211) and a second text group may be written on a second object (212). According to a comparative example, as character recognition (220) is performed on a first image (210), a character recognition result (230) may be obtained.
[0097] However, as the first text group included in the first object (211) and the second text group included in the second object (212) are placed within a relatively short distance, the electronic device (101) may fail to distinguish between the first text group and the second text group. In this case, a character recognition result (230) for a single text group including the first text group and the second text group may be obtained, as shown in FIG. 2b. For example, "on Earth, the star" in "shining on Earth, the star" written in the second document (202) may be covered by the first document (201), and accordingly, "Today, the New" of the first document (201) may be placed to the right of "shining" within the first image (210). If the distinction between the first text group and the second text group fails, the electronic device (101) may recognize the text of "shining Today, the New". For example, among the character recognition results (230) of FIG. 2b, the texts marked with an underline correspond to a first text group within the first object (211), and the remaining texts may correspond to a part of a second text group within the second object (212). As in FIG. 2b, the character recognition results (230) may include texts that are out of context due to a failure to separate the first text group and the second text group. Accordingly, a distinction between the first text group and the second text group is required based on an analysis of the physical arrangement relationship between the objects (211, 212) within the first image (210), and this will be explained below based on embodiments.
[0098] FIG. 3 is a drawing for explaining a method of operation of an electronic device according to one embodiment. The embodiment of FIG. 3 will be explained with reference to FIG. 4a to 4d. FIG. 4a to 4d are drawings for explaining the provision of analysis results associated with text.
[0099] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0100] According to one embodiment, the operations of FIG. 3 can be understood as being performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0101] According to one embodiment, the electronic device (101) may provide a first image (210), such as in FIG. 4a, in operation 301. For example, the electronic device (101) may provide (e.g., display, but not be limited to) the first image (210) stored within the electronic device (101) by running a gallery application. For example, the electronic device (101) may provide the first image (210) included in a web page by running a browsing application. For example, it may be implemented as a VR (virtual reality) device (or may be referred to as a HMD (head mounted display) or VST (video see-through) device). In this case, the electronic device (101) may display an image for VR service (e.g., an image for the left eye and an image for the right eye, but not limited to), and this may be referred to as providing the first image (210). For example, an image for a VR service may be a captured image of the real world taken by a camera, an image in which additional VR objects are placed on a captured image of the real world, or an image that is not an image of the real world (e.g., an image loaded by an electronic device (101)), but this is exemplary and not limited. For example, the electronic device (101) may be implemented as a glasses-type wearable electronic device (or may be named smart glasses), in which case a first image (210) of the real world may be obtained based on a camera included in the electronic device (101). Those skilled in the art will understand that the first image (210) may be loaded for image analysis for the purpose of providing a function (e.g., analysis of objects included in the image), and that this may be expressed as the provision of the first image (210).For example, the electronic device (101) may be implemented as a robot, in which case a first image (210) of the real world may be obtained based on a camera included in the electronic device (101). The first image (210) may be loaded for image analysis for the purpose of providing a function (e.g., analysis of objects included in the image and / or robot driving based on the analysis results), and those skilled in the art will understand that this may be expressed as providing the first image (210).
[0102] According to one embodiment, the electronic device (101) may, in operation 303, identify (or recognize) a first object (211) on which a first text group is written and a second object (212) on which a second text group is written, as shown in FIG. 4a, included in the first image (210). For example, the electronic device (101) may identify the first object (211) on which a first text group is written and the second object (212) on which a second text group is written based on a text analysis request, but those skilled in the art will understand that a text analysis request is not necessarily required. The first image (210) may include a first object (211) corresponding to a first document (201) in the real world, a second object (212) corresponding to a second document (202), and a third object (213) corresponding to a floor (203). The first object (211) may include a first text group, and the second object (212) may include a second text group. The first image (210) may be a captured image of the real world, such as in FIG. 2b, in which a first document (201) is placed on a second document (202) and the second document (202) is placed on a floor (203). For example, the electronic device (101) may identify the first object (211) and the second object (212) using an artificial intelligence model (e.g., LVM, and / or LWM, but is not limited thereto), but there is no limitation on the method of identification. For example, the electronic device (101) may perform the 303 operation using a single artificial intelligence model or perform the 303 operation using multiple artificial intelligence models, which may be applied to the operation as well as other operations described in this disclosure.For example, those skilled in the art will understand that each of the multiple operations, including at least one subsequent or prior operation and 303, may be performed using a single artificial intelligence model or using multiple artificial intelligence models. For example, an electronic device (101) may perform a specific operation based on an artificial intelligence model stored internally (which may be referred to as an on-device artificial intelligence model) or may perform a specific operation based on an artificial intelligence model stored externally (which may be referred to as a cloud artificial intelligence model or an artificial intelligence model stored on a server).
[0103] According to one embodiment, the electronic device (101) can determine, in operation 305, a physical arrangement relationship (410) between a first object (211) and a second object (212). For example, the electronic device (101) can determine a physical arrangement relationship (410) in which the first object (211) is placed on (or covers) the second object (212) in three dimensions (or in the real world). According to one embodiment, the electronic device (101) can determine, in operation 307, whether the first object (211) and the second object (212) together constitute an object in which a single text group is written. If it is determined that the objects constitute an object in which a single text group is written (operation 307—yes), according to one embodiment, the electronic device (101) can provide an analysis result associated with a third text group including the first text group and the second text group in operation 309. If it is determined that the objects each constitute the objects on which the text groups are written (Action 307—No), according to one embodiment, the electronic device (101) may provide an analysis result associated with either the first text group or the second text group in Action 311. Meanwhile, in another example, it will be understood by those skilled in the art that the electronic device (101) may provide analysis results corresponding to the first text group and the second text group, respectively, so as to be distinct.
[0104] For example, in the example of FIG. 4a, the first object (211) is placed on (or over) the second object (212), and accordingly, the electronic device (101) can confirm that the first object (211) and the second object (212) do not constitute an object with a single text group written on it, but rather that the first object (211) and the second object (212) each constitute objects with different text groups written on them. For example, the electronic device (101) may identify the first object (211) and the second object (212) and confirm that the layer corresponding to the first object (211) is a higher layer than the layer corresponding to the second object (212), but there are no limitations on the method of confirming the physical arrangement relationship (410). In this case, as in operation 311 of FIG. 3, the electronic device (101) may provide either an analysis result (430) corresponding to a first text group of a first object (211) or an analysis result (440) corresponding to a second text group of a second object (212). For example, the electronic device (101) may provide either an analysis result (430) corresponding to a first text group or an analysis result (440) corresponding to a second text group of a second object (212) depending on the user's selection. For example, the electronic device (101) may provide at least one object for selection and may provide either an analysis result (430) corresponding to a first text group or an analysis result (440) corresponding to a second text group of a second object (212) depending on user operation on the object. For example, the electronic device (101) may provide either an analysis result (430) corresponding to a first text group based on a voice command or an analysis result (440) corresponding to a second text group of a second object (212).For example, the electronic device (101) may output a voice requesting an analysis result and, based on the corresponding user voice command, may provide either an analysis result (430) corresponding to a first text group or an analysis result (440) corresponding to a second text group of a second object (212). For example, if the electronic device (101) is implemented as a glasses-type wearable device or a VR device, those skilled in the art will understand that a selection based on user gaze analysis is also possible, and there are no limitations on the method of selection by the user.
[0105] For example, the electronic device (101) may provide either an analysis result (430) corresponding to a first text group or an analysis result (440) corresponding to a second text group of a second object (212) according to text semantic analysis. For example, the electronic device (101) may provide an analysis result that satisfies specified conditions according to text semantic analysis. The specified conditions may be set based on accuracy, importance, user interest, and / or user context, but are not limited thereto. Those skilled in the art will understand that the user context may be analyzed based, for example, the analysis result of the first image (210) and / or images acquired before and / or after the first image (210).
[0106] For example, if the electronic device (101) is implemented as a robot, the analysis result may be selected according to the mission assigned to the electronic device (101). For example, if the electronic device (101) is assigned a mission to analyze an upper document, the electronic device (101) may provide an analysis result (430) for the first object (211). For example, if the electronic device (101) is assigned a mission to analyze a lower document, the electronic device (101) may provide an analysis result (440) for the second object (212). For example, if the electronic device (101) is assigned a mission to analyze the entire document, the electronic device (101) may provide an analysis result (430) for the first object (211) and an analysis result (440) for the second object (212). In this case, after providing the analysis result (430) for the first object (211), the robot may separate and remove the first document (201) corresponding to the first object (211) from the second document (202). Afterward, the robot may photograph the second document (202) and provide the analysis result (440). Meanwhile, as this is exemplary, the electronic device (101) may provide the analysis result (440) for the second object (212) using only the first image (210), which will be described later.
[0107] Referring again to FIG. 4a, when the electronic device (101) provides an analysis result (440) corresponding to the second object (212), a part (441) of the analysis result (440) may not be obtainable from the first image (210). The part (441) of the analysis result (440) corresponds to a part covered by the first document (201) in the real world, and accordingly, the first image (210) does not contain the corresponding text. The electronic device (101) may generate text corresponding to the part (441) based on the preceding and succeeding context, for example. For example, the part (441) based on the preceding and succeeding context may be generated based on LLM, but there are no limitations. For example, the electronic device (101) may provide the analysis result (440) by verifying the part (441) based on a web search (or external database search), for example, based on a RAG technique. For example, the full text of the second document (202) corresponding to the second object (212) can be verified through a web search of at least a portion of the second text group corresponding to the second object (212). The electronic device (101) can provide the full text obtained through the web search as an analysis result (440) for the second text group, and accordingly, a portion (441) that was not verified due to the obscuration of the first object (211) may be provided. For example, the electronic device (101) may automatically perform the generation of the portion (441) that was not verified due to obscuration, or may perform it based on a user command, and this will be described later.
[0108] Referring to FIG. 4b, the electronic device (101) may provide a first image (240). The electronic device (101) may identify a first object (241), a second object (242), and a background object (243) corresponding to the background from the first image (240). The electronic device (101) may identify, for example, a first object (241) on which a first text group is written and a second object (242) on which a second text group is written. As described with reference to FIG. 3, the electronic device (101) may determine whether the first object (241) and the second object (242) are objects on which a single text group is written. For example, the electronic device (101) can determine whether the first object (241) and the second object (242) are objects on which a single text group is written, based on confirming that the layer corresponding to the first object (241) and the layer corresponding to the second object (242) are the same layer. For example, the electronic device (101) can determine the layer corresponding to the first object (241) based on the physical relationship between the first object (241) and the background object (243), and / or determine the layer corresponding to the first object (241) based on the physical relationship between the first object (241) and the background object (243), but this is exemplary and there are no limitations on the method of determining the layer. In the example of FIG. 4b, for instance, the layer corresponding to the first object (241) identified based on the physical relationship of the first object (241) to the background object (243) and the layer corresponding to the first object (241) identified based on the physical relationship of the first object (241) to the background object (243) can be identified as being the same. Accordingly, the electronic device (101) can be identified that the first object (241) and the second object (242) constitute a single object on which a single text group is written.For example, the electronic device (101) can confirm that the shape of a part (241a) of the first object (241) and the shape of a part (242a) of the second object (242) correspond to each other, and accordingly, the first object (241) and the second object (242) can be confirmed to constitute a single object on which a single text group is written. In the example of FIG. 4b, the shape of a part (241a) of the first object (241) and the shape of a part (242a) of the second object (242) can be confirmed to correspond to shapes of a torn document, but there are no limitations. For example, the electronic device (101) can confirm that the material of the first object (241) and the material of the second object (242) are substantially the same, and accordingly, the first object (241) and the second object (242) can be confirmed to constitute a single object on which a single text group is written. For example, the electronic device (101) can determine the material of the objects (241, 242) based on LVM or LWM, but this is exemplary and not limited. For example, the electronic device (101) can determine that a first text group written on the first object (241) and a second text group written on the second object (242) constitute a single text group, and accordingly, the first object (241) and the second object (242) can be determined to constitute a single object on which a single text group is written. For example, in the example of FIG. 4b, the electronic device (101) can see that “Once upon a ti” written on the first object (241) and “me a little star” written on the second object (242) form part of a sentence, and thus the first object (241) and the second object (242) form one object with a group of text written on it, but there are no limitations.The electronic device (101) can determine the physical arrangement relationship (412) of the objects (241, 242) based on at least some of the various methods described above. Based on the physical arrangement relationship (412) of the objects (241, 242), the electronic device (101) can determine that the objects (241, 242) constitute a single object on which a single text group is written. Accordingly, the electronic device (101) can perform character recognition (420), which is an example of an analysis result for a single text group corresponding to the objects (241, 242), and accordingly, an analysis result (450) for a single text group can be provided.
[0109] Referring to FIG. 4c, the electronic device (101) may provide a first image (250). The electronic device (101) may identify a first object (251), a second object (252), a third object (253), and a background object (254) corresponding to the background from the first image (250). The electronic device (101) may identify a physical arrangement relationship (413) between the objects (251, 252, 253, 254). The electronic device (101) may identify, for example, that the third object (253) is placed on the first object (251) and the second object (252). For example, the electronic device (101) may identify that the first object (251) and the second object (252) are placed on the background object (254). For example, the electronic device (101) can determine that, for example, the first object (251) and the second object (252) are placed on the same layer. The electronic device (101) can determine, based on the physical arrangement relationship (413), that the first object (251) and the second object (252) constitute an object on which a text group is written, and that the third object (253) is placed on the first object (251) and the second object (252). For example, the electronic device (101) can determine whether to perform an analysis (e.g., character recognition (420)) associated with the text group corresponding to the third object (253), or to perform an analysis (e.g., character recognition (420)) associated with the text group corresponding to the first object (251) and the second object (252). Accordingly, the electronic device (101) can provide an analysis result (460) associated with a text group corresponding to the third object (253), as in FIG. 4c, or an analysis result (470) associated with a text group corresponding to the first object (251) and the second object (252).As described above, the text (471) corresponding to the part obscured by the third object (253) may be automatically generated (e.g., generated based on LLM or generated based on RAG), but there are no limitations. Meanwhile, as described above, the electronic device (101) may perform both analysis associated with the text group corresponding to the third object (253) (e.g., character recognition (420)) and analysis associated with the text group corresponding to the first object (251) and the second object (252) (e.g., character recognition (420)).
[0110] Referring to FIG. 4d, the electronic device (101) may provide a first image (260). The electronic device (101) may identify a first object (261), a second object (262), and a background object (263) corresponding to the background from the first image (260). The electronic device (101) may identify a physical arrangement relationship (414) between the objects (261, 262, 263). The electronic device (101) may identify, for example, that the first object (261) is placed on the second object (252) and the second object (262) is placed on the third object (263). For example, the electronic device (101) may identify, for example, that the first object (261) and the second object (262) are placed on different layers. The electronic device (101) can determine, based on the physical arrangement relationship (414), that the first object (261) and the second object (262) do not constitute an object on which a single text group is written. The electronic device (101) can determine that the first object (261) does not contain (or has no text group written on it) and, accordingly, perform an analysis (e.g., character recognition (420)) associated with the text group corresponding to the second object (262). The electronic device (101) can provide an analysis result (440) associated with the text group corresponding to the second object (262), as shown in FIG. 4d. As described above, the text (443) corresponding to the part obscured by the first object (261) may be automatically generated (e.g., generated based on LLM or generated based on RAG), but there are no limitations.
[0111] FIG. 4e is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0112] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0113] According to one embodiment, the operations of FIG. 4e can be understood as being performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0114] According to one embodiment, the electronic device (101) may recognize an object corresponding to a text area from the first image in operation 471. For example, as described above, the electronic device (101) may display (or represent) the first image through the execution of an application (e.g., a gallery application or a browsing application, but is not limited thereto), but is not limited thereto. Alternatively, those skilled in the art will understand that the first image may not be displayed (or represent) if the electronic device (101) is implemented as a glasses-type wearable electronic device or a robot (or may be otherwise, but is not limited thereto). The electronic device (101) may determine context based on objects contained in the first image, for example, using an artificial intelligence model (e.g., LVM, but is not limited thereto). For example, the electronic device (101) may recognize a unit containing text as one object and recognize other areas as a background area, but is not limited thereto. For example, in the example of FIG. 4b, objects (241, 242) corresponding to each of the torn documents are recognized as objects corresponding to text, and other background areas (243) may be recognized. In operation 473, the electronic device (101) may perform text extraction and / or image analysis on the recognized objects. For example, the electronic device (101) may extract pure text data from the text area, in which case an OCR and / or artificial intelligence model (e.g., LVM, but not limited thereto) may be used, but not limited thereto. According to one embodiment, in operation 471, the electronic device (101) may perform text extraction on an extracted object-by-object basis, but not limited thereto. For example, when extracting text, text characteristics (e.g., font type, and / or text size) and / or document characteristics (e.g., document format) may be used.For example, the electronic device (101) may perform image analysis using LWM. For example, the type (or material) of the text area may be analyzed, but there are no limitations. According to one embodiment, the electronic device (101) may identify the physical arrangement relationship between objects corresponding to the text area in operation 475. For example, the electronic device (101) may identify the physical arrangement relationship between objects corresponding to the text area based on text recognition results and / or image analysis results using LWM. For example, the electronic device (101) may identify foreground objects and background objects. A foreground object may refer to an object placed relatively higher in terms of three-dimensional depth, and a background object may refer to an object placed relatively lower. For example, in the embodiment of FIG. 4b, the first object (241) and the second object (242) may be identified as foreground objects, and the background object (243) may be identified as a background object. For example, in the embodiment of FIG. 4a, when the first object (211) is a foreground object, the second object (212) may be identified as a background object, and when the second object (212) is a foreground object, the background object (213) may be a background object. Based on the identification of the foreground object and background object described above, layers for each object may be identified. For example, LWM may be specialized for understanding the real world, and accordingly, spatial characteristics between objects may be identified, and thus physical arrangement relationships between objects may be identified. For example, an electronic device (101) may configure layers for each object using LWM, and physical arrangement relationships between objects may be identified according to the analysis of the layers.For example, the electronic device (101) can sequentially check from the first layer to the last layer, thereby checking the physical arrangement relationship between each of the plurality of objects. For example, the electronic device (101) can check objects placed on the same layer. In this case, the shape (contact, angle, and / or length), character recognition result, and / or material of each of the objects may be used. According to one embodiment, the electronic device (101) can generate at least one group corresponding to each of the objects in the 477 operation. For example, objects placed on each of different layers may correspond to different groups, and objects placed on one layer that satisfy different conditions (e.g., conditions based on shape, character recognition result, and / or material) may correspond to one group, but are not limited thereto. As described above, the electronic device (101) may provide analysis results by text group, and may provide reconstructed objects based on objects determined to be a single group, which will be described later. For example, the electronic device (101) may provide information on detailed groups based on context analysis by group (e.g., may be representative keywords, but is not limited thereto), which will be described later. For example, the electronic device (101) may provide relationships between keywords and groups, and may provide group control functions based on user operation, which will be described later.
[0115] FIG. 5 is a drawing for explaining a method of operation of an electronic device according to one embodiment. The embodiment of FIG. 5 will be described with reference to FIG. 6a to 6d. FIG. 6a to 6d are examples of screens displayed by an electronic device according to various embodiments.
[0116] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0117] According to one embodiment, the operations of FIG. 5 can be understood as being performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0118] According to one embodiment, the electronic device (101) may provide a first image (240) in operation 501, for example, as in FIG. 6a. According to one embodiment, the electronic device (101) may identify a first object (241) on which a first text group is written and a second object (242) on which a second text group is written in operation 503. According to one embodiment, the electronic device (101) may identify at least a portion of the first text group or the second text group as text to be recognized in operation 505, based on the physical arrangement relationship between the first object (241) and the second object (242). For example, the electronic device (101) may identify the entire first text group or the second text group as text to be recognized. For example, the electronic device (101) can identify the first text group and the second text group as the entire text to be recognized by identifying the first object (241) and the second object (242) as constituting an object on which a text group is written. According to one embodiment, the electronic device (101) can provide, in operation 507, an analysis result associated with the text to be recognized and a second image (610) including a third object (611) generated based on adjustments to at least some of the first object (241) and the second object (242). In the example of FIG. 6a, the electronic device (101) can, for example, identify that the first object (241) and the second object (242) correspond to torn parts of a document, and accordingly provide a third object (611) corresponding to the original document before it was torn. In this case, the electronic device (101) may provide a third object (611) based on the movement and / or deletion of at least one of the first object (241) and the second object (242), but this is exemplary and there is no limitation on the method of adjustment.For example, in the embodiment of FIG. 6a, affordances (614, 615, 616, 617) for selecting an analysis result are disclosed, but this is exemplary and there is no limit to the type and / or number of affordances. Those skilled in the art will understand that an affordance is an element that causes the performance of a preset function upon selection (or designation), for example, and may be named an icon, a visual element, or an element, a visual object, or an object. Alternatively, those skilled in the art will understand that an analysis result may be provided without providing an affordance.
[0119] Referring to FIG. 6b, the electronic device (101) may provide a first image (210). The first image (210) may include a first object (211) corresponding to a first document (201) in the real world, a second object (212) corresponding to a second document (202), and a third object (213) corresponding to a floor (203). As described above, the electronic device (101) may provide an analysis result for either a first text group corresponding to the first object (211) or a second text group corresponding to the second object (212), based on confirming that each of the first object (211) and the second object (212) corresponds to each of the objects on which each of the different text groups is written. In the embodiment of FIG. 6b, it is assumed that the electronic device (101) identifies a second text group written on a second object (212) as an object for analysis. The electronic device (101) may generate text (622) corresponding to the part obscured by the first object (211) based on LLM or RAG as described above. The electronic device (101) may provide an image (620) containing a fourth object (621) (which may be referred to as a reconstructed object) on which a text group containing the generated text (622) is written. The generation of the fourth object (621) may be based on the adjustment of the first object (211) and / or the second object (212). For example, a fourth object (621) may be created by removing (or, may be referred to as deletion or movement) the first object (211) of the first image (210), generating text based on LLM and / or RAG, and / or creating a part (622) based on the generated text (for example, generating text of a font and size corresponding to a second text group on a material corresponding to the second object (212).
[0120] Referring to FIG. 6c, the electronic device (101) may provide a first image (210). The electronic device (101) may provide an analysis result for either a first text group corresponding to the first object (211) or a second text group corresponding to the second object (212), based on confirming that each of the first object (211) and the second object (212) corresponds to each of the objects on which each of the different text groups is written. In the embodiment of FIG. 6c, it is assumed that the electronic device (101) confirms the second text group written on the second object (212) as the subject of analysis. In this case, the electronic device (101) may provide an image (630) including an object (631) that visually distinguishes from the surroundings a part (633) that is obscured by the first object (211) and is difficult to verify text from, as in FIG. 6c. The electronic device (101) may provide affordances (635) that cause automatic generation. When a selection for affordances (635) is confirmed, the electronic device (101) may provide a fourth object (621), such as in FIG. 6b, for example. Meanwhile, the method of selecting affordances (635) is merely exemplary, and a fourth object (621), such as in FIG. 6b, may be provided based on a specified user command for a part (633) (e.g., a rubbing activity, but is not limited to). In this case, in one example, a graphic effect that changes to the generated content for the part that the user rubs may be applied, but this is merely exemplary.
[0121] Referring to FIG. 6d, the electronic device (101) may provide a first image (250). The first image (250) may include a first object (251) and a second object (252) corresponding to a torn document, a third object (253) corresponding to a document placed on the torn document, and a background object (254). For example, the electronic device (101) may determine that the first object (251) and the second object (252) constitute a single object on which a single group of text is written. For example, the electronic device (101) may determine that the third object (253) constitutes objects on which different groups of text are written, respectively from the first object (251) and the second object (252). In the example of FIG. 6d, it is assumed that the first text group is selected from among the first text group corresponding to the first object (251) and the second object (252) and the second text group corresponding to the third object (253). The electronic device (101) can generate text (622) corresponding to the part obscured by the third object (253) based on LLM or RAG as described above. The electronic device (101) can provide an image (620) containing a fourth object (621) (which may be referred to as a reconstructed object) on which a text group containing the generated text (622) is written. The generation of the fourth object (621) can be generated based on the adjustment of the first object (251), the second object (252), and / or the third object (253).For example, a fourth object (621) may be created by removing (or, may be referred to as deletion or movement) a third object (253) of a first image (250), generating text based on LLM and / or RAG, and / or creating a part (622) based on the generated text (e.g., generating text of a font and size corresponding to a first text group on a material corresponding to objects (251, 252), and / or combining according to the movement of objects (251, 252). For example, similar to what is described in FIG. 6c, the electronic device (101) may provide the fourth object (621) based on one or more user inputs, and at least one intermediate restored image may be provided. For example, the electronic device (101) may provide an intermediate restored image (e.g., the first image (240) of FIG. 6a) from which the third object (253) has been removed, based on a user input (e.g., a flick gesture) that causes the removal of the third object (253), and then provide a fourth object (621) based on a user input (e.g., a pinch-in gesture) that causes the combination of both objects (241, 242).
[0122] FIG. 7 is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0123] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0124] According to one embodiment, the operations of FIG. 7 can be understood as being performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0125] According to one embodiment, the electronic device (101) may provide, in operation 701, a first image comprising a first object on which a first text group is written and a second object on which a second text group is written. According to one embodiment, the electronic device (101) may, in operation 703, provide a second image comprising a third object created based on adjustments to at least some of the first object and the second object and a third text group written on the third object, and an analysis result associated with the third text group. Here, adjustments to at least some of the first object and the second object may include moving the object, removing the object, changing the shape of the object, and / or changing the material of the object, but this is exemplary and there is no limitation on the method of adjustment. For example, the third object may be created based on additional content creation (e.g., text creation). Here, the third text group may be created based on at least some of the first text group or the second text group. For example, the third text group may be at least part of the combined result of the first text group and the second text group. For example, the third text group may be either the first text group or the second text group.
[0126] FIG. 8 is a drawing for explaining a method of operation of an electronic device according to one embodiment. The embodiment of FIG. 8 will be explained with reference to FIG. 9. FIG. 9 is a drawing for explaining a classification result according to one embodiment.
[0127] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0128] According to one embodiment, the operations of FIG. 8 can be understood as being performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0129] According to one embodiment, the electronic device (101) may provide a first image (240) as in FIG. 9 in operation 801. The first image (240) may include a first object (241) and a second object (242) corresponding to a torn document. According to one embodiment, the electronic device (101) may identify a first object (241) on which a first text group is written and a second object (242) on which a second text group is written in operation 803. According to one embodiment, the electronic device (101) may identify at least a portion of the first text group or the second text group as text to be recognized in operation 805, based on the physical arrangement relationship (410) between the first object (241) and the second object (242). For example, the electronic device (101) can identify at least a portion of the first text group and the second text group as text to be recognized by determining, based on the physical arrangement relationship (410) of the first object (241) and the second object (242), that the first object (241) and the second object (242) constitute an object on which a text group is written. The electronic device (101) can perform, for example, character recognition (420) on the text to be recognized. According to one embodiment, the electronic device (101) can provide a classification result for at least one portion of the third object and a second image including a third object generated based on the adjustment of at least a portion of the first object (241) and the second object (242) in operation 807. For example, referring to FIG. 9, visual elements (911, 912, 913, 914) representing the classification result may be provided. For example, the electronic device (101) can classify the document by paragraphs and may be provided with visual elements (911, 912, 913, 914) that allow the classified parts to be visually identified.Classification by paragraph may be based on formal conditions such as, for example, content, amount of text, and / or line breaks, but those skilled in the art will understand that there are no restrictions on the method of classification.
[0130] FIG. 10 is a diagram illustrating a method of operation of an electronic device according to one embodiment. The embodiment of FIG. 10 will be described with reference to FIG. 11a. FIG. 11a is a diagram illustrating the determination of a text to be recognized according to a user selection according to one embodiment.
[0131] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0132] According to one embodiment, the operations of FIG. 10 can be understood as being performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0133] According to one embodiment, the electronic device (101) may provide a first image (210) in operation 1001, for example, as in FIG. 11a. According to one embodiment, the electronic device (101) may identify a plurality of objects (211, 212) in operation 1003. According to one embodiment, the electronic device (101) may identify a physical arrangement relationship between the plurality of objects (211, 212) in operation 1005. According to one embodiment, the electronic device (101) may identify a plurality of recognizable texts based on the physical arrangement relationship in operation 1007. For example, in the embodiment of FIG. 11a, a first recognition target text (1101), a second recognition target text (1111), a third recognition target text (1112), a fourth recognition target text (1113), a fifth recognition target text (1114), and a sixth recognition target text (1120) may be identified. According to one embodiment, the electronic device (101) may confirm a selection of the first recognition target text through a user interface that allows at least some of the plurality of recognition target texts to be selected in operation 1009. According to one embodiment, the electronic device (101) may provide an analysis result for the selected first recognition target text in operation 1011. In the example of FIG. 11a, the user interface may include a progress bar (1134) and a selector (or indicator, which may be named) (1135) placed (or pointing) to a position on the progress bar (1134), but this is exemplary. Visual elements (1131, 1132, 1133) associated with texts to be recognized (1101, 1111, 1112, 1113, 1114, 1120) may be provided to correspond to at least one position on the progress bar (1134).For example, when a selector (1135) is placed at a first position of a progress bar (1134), a visual element (1131) associated with the first position may be displayed to be distinguishable from other visual elements (1132, 1133). The first visual element (1131) may be associated, for example, with a first text to be recognized (1101) and may include information associated with the first text to be recognized (1101) (here, a keyword called stock). Based on the operation of the selector (1135) and / or additional user confirmation (for example, selection of the first text to be recognized (1101)), the first text to be recognized (1101) may be selected. For example, when a selector (1135) is placed at a second position of a progress bar (1134), a visual element (1132) associated with the second position may be represented to be distinguishable from other visual elements (1131, 1133). The second visual element (1132) is associated with, for example, a second text to be recognized (1111), a third text to be recognized (1112), a fourth text to be recognized (1113), and a fifth text to be recognized (1114), and may include classification results and information (here, a keyword called 'para' indicating classification by paragraph). Based on the operation of the selector (1135) and / or additional user confirmation (e.g., selection of the texts to be recognized (1111, 1112, 1113, 1114)), at least some of the texts to be recognized (1111, 1112, 1113, 1114) may be selected. For example, when the selector (1135) is placed at a third position of the progress bar (1134), the visual element (1133) associated with the third position may be displayed to be distinguishable from other visual elements (1131, 1132). The third visual element (1133) may be associated with, for example, the sixth text to be recognized (1120) and may include information associated with the sixth text to be recognized (here, the keyword 'story').Based on the operation of the selector (1135) and / or additional user confirmation (e.g., selection of the sixth recognizable text (1120)), the sixth recognizable text (1120) may be selected. Meanwhile, those skilled in the art will understand that the user interface described above is merely exemplary and there are no restrictions on its implementation.
[0134] FIG. 11b is a drawing for illustrating an image provided by an electronic device according to one embodiment.
[0135] According to one embodiment, the electronic device (101) may further display a progress bar (1164) and a selector (1165) on a first image (210) comprising a first object (241), a second object (242), and a third object (244). In the example of FIG. 11b, the electronic device (101) may provide an image restoration process in response to the operation of the selector (1165). For example, when the selector (1165) is placed at a first position of the progress bar (1164), the electronic device (101) may provide an original image. In this case, the text to be recognized (1171) may be selected through additional user input or directly without additional user input. For example, when the selector (1165) is positioned at a second position of the progress bar (1164), the electronic device (101) may provide an intermediate restored image in which the third object (244) is excluded and content created in place of the area where the third object (244) was located is included. For example, when the selector (1165) is positioned at a third position of the progress bar (1164), the electronic device (101) may provide a restored image in which the first object (241) and the second object (242) are combined with each other. In this case, the text to be recognized (1172) may be selected through additional user input or directly without additional user input.
[0136] FIG. 12a is a drawing for explaining a method of operation of an electronic device according to one embodiment. The embodiment of FIG. 12a will be explained with reference to FIG. 12b. FIG. 12b is a drawing for explaining an image provided by an electronic device according to one embodiment.
[0137] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0138] According to one embodiment, the operations of FIG. 12a can be understood as being performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0139] According to one embodiment, the electronic device (101) may provide a first image (240) as shown in FIG. 12b in operation 1201. According to one embodiment, the electronic device (101) may identify a first object (241) on which a first text group is written and a second object (242) on which a second text group is written in operation 1203. According to one embodiment, the electronic device (101) may identify at least a portion of the first text group or the second text group as text to be recognized in operation 1205, based on the physical arrangement relationship between the first object (241) and the second object (242). In FIG. 12b, it is assumed that the first text group and the second text group are identified as text to be recognized. According to one embodiment, as described above, the electronic device (101) may provide additional information (1214) associated with the text to be recognized and a second image (620) comprising a third object (object (621) in FIG. 12b) generated based on adjustments to at least some of the first object (241) and the second object (242) in operation 1207. For example, in the example of FIG. 12b, the additional information (1214) may cause the addition of a user signature. For example, when a connection based on user information is completed, the electronic device (101) may provide additional information (1214) associated with the user information and the text to be recognized. If a selection for the additional information (1214) is confirmed, the electronic device (101) may add a signature previously stored associated with a user account, for example, to the second image (620) (or the result of the analysis thereof). Data associated with the signature may be provided by other applications (e.g., applications associated with authentication, but without limitation), but without limitation.For example, the electronic device (101) can confirm that the document requires a signature based on the character recognition result of the third object (object (621) in FIG. 12b), and thus confirm that a user signature is required. Meanwhile, the addition of the signature is merely illustrative, and there are no restrictions on the method of implementing the additional information.
[0140] FIG. 13 is a drawing for explaining a method of operation of an electronic device according to one embodiment. The embodiment of FIG. 13 will be explained with reference to FIG. 14a to 14c. FIG. 14a to 14c is a drawing for explaining an image provided by an electronic device according to one embodiment.
[0141] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0142] According to one embodiment, the operations of FIG. 13 can be understood as being performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0143] According to one embodiment, the electronic device (101) can identify a plurality of images in operation 1301. The plurality of images may be, for example, images taken over time, or frames that constitute a video. For example, the electronic device (101) can identify the first image (1410) in FIG. 14a, the intermediate image (1420) in FIG. 14b, and the second image (1430) in FIG. 14c. As described above, the images (1410, 1420, 1430) may be, for example, images that constitute a video, but are not limited thereto. According to one embodiment, the electronic device (101) can identify a first object on which a first text group is written, which is included in the first image (1410) among the plurality of images, in operation 1303. According to one embodiment, the electronic device (101) can identify a second object on which a second text group included in a second image (1430) among a plurality of images is written in operation 1305. According to one embodiment, the electronic device (101) can identify the temporal and spatial relationship between the first object and the second object based on the first image (1410), the second image (1430), and / or at least one intermediate image (1420) in operation 1307. According to one embodiment, the electronic device (101) can identify at least some of the first text group and the second text group as text to be recognized in operation 1309. According to one embodiment, the electronic device (101) can provide an analysis result associated with the text to be recognized in operation 1311. For example, referring to FIG. 14a, the first image (1410) may include a background object (1411) and an object (1412) corresponding to a book.The object (1412) corresponding to the book may include, for example, a sub-object (1412a) corresponding to the first page (Page #1) of the book and a sub-object (1412b) corresponding to the second page (Page #2) of the book. For example, referring to FIG. 14b, the intermediate image (1420) may include a background object (1421) and an object (1422) corresponding to the book. The object (1422) corresponding to the book may include, for example, a sub-object (1422a) corresponding to the first page (Page #1) of the book, a sub-object (1422b) corresponding to the second page (Page #2) being turned over, and a sub-object (1422c) corresponding to a part of the fourth page (Page #4). For example, referring to FIG. 14c, the second image (1430) may include a background object (1431) and an object (1432) corresponding to the book. The object (1432) corresponding to the book may include, for example, a sub-object (1432a) corresponding to the third page (Page #3) of the book and a sub-object (1432b) corresponding to the fourth page (Page #4). The electronic device (101) may perform an analysis of the images (1410, 1420, 1430) (e.g., an analysis based on LVM and / or LWM, but is not limited thereto). Based on the results of the analysis, the electronic device (101) may recognize that the images (1410, 1420, 1430) relate to a situation where the first page (Page #1) and the second page (Page #2) of the book are opened, and then the third page (Page #3) and the fourth page (Page #4) are opened as the pages are turned. Accordingly, it can be confirmed that the temporal and spatial relationship of the sub-object (1422b) corresponding to the second page (Page #2) of the first image (1410) and the sub-object (1432a) corresponding to the third page (Page #3) of the second image (1430) is the relationship of consecutive pages of a book.Accordingly, the electronic device (101) can confirm that the text group written in the sub-object (1422b) corresponding to the second page (Page #2) and the text group corresponding to the sub-object (1432a) corresponding to the third page (Page #3) constitute a single text group. For example, regarding the sentence “Once upon a time, a little star twinkled alone in the vast night sky.”, “Once upon a time, a little star” may be written on the second page (Page #2), and “twinkled alone in the vast night sky.” may be written on the third page (Page #3). The electronic device (101) can recognize the complete sentence "Once upon a time, a little star twinkled alone in the vast night sky." based on the temporal and spatial relationship between the sub-object (1422b) corresponding to the second page (Page #2) and the sub-object (1432a) corresponding to the third page (Page #3), rather than recognizing "Once upon a time, a little star" and "twinkled alone in the vast night sky." by recognizing characters by separating the first image (1410) and the third image (1430).
[0144] FIG. 15 is a drawing for explaining a method of operation of an electronic device according to one embodiment. The embodiment of FIG. 15 will be described with reference to FIG. 16a to FIG. 16c. FIG. 16a to FIG. 16c are drawings for explaining text analysis based on objects not associated with text.
[0145] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0146] According to one embodiment, the operations of FIG. 15 can be understood as being performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0147] According to one embodiment, the electronic device (101) may provide a first image in operation 1501. As described above, when the electronic device (101) is implemented as a glasses-type wearable electronic device, it has been explained that the acquisition and / or loading of an image corresponding to the actual environment acquired by a camera may be the provision of the image. The electronic device (101) may be implemented, for example, as a glasses-type wearable electronic device, and the first image (1610) of FIG. 16a may be, for example, an image of the actual environment captured by a camera included in the glasses-type wearable electronic device, but is not limited thereto. According to one embodiment, the electronic device (101) may identify a first object (1641) with a first text group written on it and a second object (1612) with a second text group written on it, which are included in the first image (1610), in operation 1503. The first object (1641) may be recognized as being displayed on the screen (1640) of, for example, an external electronic device (e.g., a smartphone) (e.g., the electronic device (102, 104) of FIG. 1). The first object (1641) may be, for example, an object corresponding to a text message. The second object (1612) may be, for example, a menu board. According to one embodiment, the electronic device (101) may, in operation 1505, confirm the physical arrangement relationship between the first object (1641) and the second object (1612). According to one embodiment, the electronic device (101) may, in operation 1507, confirm that the first object (1641) and the second object (1612) are objects corresponding to different text groups. For example, as described above, based on the layers of the first object (1641) and the second object (1612), it may be confirmed that the first object (1641) and the second object (1612) are objects corresponding to different text groups, but there are no limitations.According to one embodiment, the electronic device (101) may, in operation 1509, select either a first text group or a second text group based on a third object (1613) included in the first image (1610). For example, the third object (1613) in FIG. 16a may correspond to the environment of a restaurant. The electronic device (101) may select, for example, a second object (1612) based on a context analysis result based on a third object (1613) without text and / or at least one object (1641, 1612) with text. The context may be analyzed, for example, as "a waiter is approaching from the restaurant and a menu selection is required." Accordingly, a second object (1612) corresponding to the context analysis result may be selected, but this is exemplary and not limited. For example, the first object (1641) may be identified as an object corresponding to a text message and not associated with the current context, which is a menu selection. According to one embodiment, the electronic device (101) may provide an analysis result associated with a selected text group in the 1511 operation. For example, the electronic device (101) may identify that at least a portion of the second object (1612) is obscured by the smartphone. The electronic device (101) may provide the obscured text portion to the user. For example, the electronic device (101) may identify information associated with a restaurant menu based on a search through a browser. For example, as shown in FIG. 16b, the electronic device (101) may control the smartphone being used by the user to provide information (1611) associated with a menu including at least one menu (1611a, 1611b, 1611c).For example, in the example of FIG. 16c, the electronic device (101) (e.g., a glasses-type wearable electronic device) may provide a voice (1630) inquiring whether to recommend a representative menu, either directly or through another accessory (e.g., a wireless earphone or an external electronic device, but not limited thereto). If a request for a representative menu recommendation is received from a user, the electronic device (101) may provide a command to an external electronic device (smartphone) that triggers a representative menu recommendation. The external electronic device (smartphone) may provide content associated with the representative menu recommendation in response to the reception of the command. For example, the external electronic device (smartphone) may provide content based on obtaining user preference information for food through a gallery application based on an AI assistant, and / or based on searching for the restaurant's representative menu through an internet browser search, but not limited thereto. Content may be generated, for example, based on data (e.g., menus and / or descriptions) additionally provided by the electronic device (101).
[0148] FIG. 17 is a drawing for explaining a method of operation of an electronic device according to one embodiment. The embodiment of FIG. 17 will be explained with reference to FIG. 18. FIG. 18a and FIG. 18b are drawings for explaining text analysis.
[0149] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0150] According to one embodiment, the operations of FIG. 17 can be understood as being performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0151] According to one embodiment, the electronic device (101) may provide a first image in operation 17501. As described above, when the electronic device (101) is implemented as a glasses-type wearable electronic device, it has been explained that the acquisition and / or loading of an image corresponding to the actual environment acquired by a camera may be the provision of the image. The electronic device (101) may be implemented, for example, as a glasses-type wearable electronic device, and the first image (1810) of FIG. 18a may be, for example, an image of the actual environment captured by a camera included in the glasses-type wearable electronic device, but is not limited thereto. According to one embodiment, the electronic device (101) may identify a first object (1811) with a first text group written on it and a second object (1812) with a second text group written on it, which are included in the first image (1810), in operation 1703. According to one embodiment, the electronic device (101) can determine the physical arrangement relationship between the first object (1811) and the second object (1812) in operation 1705. According to one embodiment, the electronic device (101) can determine in operation 1707 that the first object (1811) and the second object (1812) are objects corresponding to different text groups, respectively. For example, as described above, it may be determined that the first object (1811) and the second object (1812) are objects corresponding to different text groups, respectively, based on the layers of the first object (1811) and the second object (1812), but is not limited thereto. According to one embodiment, the electronic device (101) can control the second object (1812) to reflect the analysis result associated with the first text group in operation 1709. For example, the electronic device (101) can analyze the context based on at least one object included in the first image (1810).For example, in the embodiment of FIG. 18a, the context of "taking notes on learning content using a tablet PC while watching a presentation in a classroom" can be analyzed. Based on the result of the context analysis, the electronic device (101) can confirm that the first object (1811) is a reference object. For example, if synchronization of handwriting is required, other content may be updated according to the content of one of the contents, and the object serving as the standard for synchronization may be named the reference object, but there are no limitations. The electronic device (101) can compare the first text group written on the first object (1811) and the second text group written on the second object (1812). Based on the result of the comparison, it can be confirmed that at least a portion of the first text group (for example, additional handwriting by a professor, but there are no limitations) is not included in the second text group. The electronic device (101) may control at least a portion of a first text group not included in a second text group to be reflected in a second object (1812). In one example, the electronic device (101) may provide data that causes the reflection of at least a portion of a first text group not included in a second text group to an external electronic device (e.g., a tablet) corresponding to the second object (1812). The external electronic device (e.g., a tablet) (e.g., the electronic device (102, 104) of FIG. 1) may reflect at least a portion of a first text group not included in a second text group onto an image displayed on a screen in response to receiving data, but those skilled in the art will understand that this is exemplary and there are no limitations on the method of reflection. Alternatively, the electronic device (101) may generate data to explain the difference and provide it to a storage (or external electronic device (e.g., a tablet)).Alternatively, the electronic device (101) may confirm that the slide (or page) corresponding to the second text group and the slide (or page) corresponding to the first text group are different, and may provide data to an external electronic device (e.g., a tablet) that causes a screen transition to the slide (or page) corresponding to the second text group. The external electronic device (e.g., a tablet) may display the slide (or page) corresponding to the first text group in response to receiving the data.
[0152] Referring to FIG. 18b, the electronic device (101) can identify a first object (1821) on which a first text group is written and a second object (1822) on which a second text group is written, which are included in the first image (1820). The electronic device (101) may store and / or provide data to explain the difference, for example, based on the difference between the first text group and the second text group. For example, the electronic device (101) may store data to explain the difference in conjunction with the time when the first image (1820) was taken, and may store and / or provide data to explain the difference at multiple times based on other images. Those skilled in the art will understand that the electronic device (101) may generate data for explanation, for example, using voice data (for example, data corresponding to the lecturer's voice).
[0153] FIG. 19 is a drawing for explaining the operation of an electronic device according to one embodiment.
[0154] In the example of FIG. 19, the electronic device (101) may be implemented as a robot. For example, the electronic device (101) may use a camera to take an image of an environment, such as a courier warehouse. The image may include box objects (1930, 1940). The image may include invoice objects (1931, 1932) and an invoice object (1941).
[0155] For example, the electronic device (101) can identify invoice objects (1931, 1932) and invoice object (1941) from an image. The electronic device (101) can identify the physical arrangement relationship between the invoice objects (1931, 1932) and invoice object (1941) and / or box objects (1930, 1940), for example, using LVM and / or LWM. The electronic device (101) can, for example, recognize areas where boxes are clustered and empty spaces separately. The electronic device (101) can, although not illustrated, recognize a text area outside the box, an invoice area, and / or a shelf number area separately. The electronic device (101) can recognize each of the text areas corresponding to the box, invoice, and / or shelf number, and can, for example, use an artificial intelligence model to distinguish the box, invoice sticker, and shelf according to the context of the image. The electronic device (101) may identify invoice objects (1931, 1932) and invoice object (1941) as foreground objects and box objects (1930, 1940) (or objects corresponding to shelves, not shown) as background objects based on physical placement relationships between objects, for example using LVM and / or LWM, and the physical placement relationships may be identified based on the layers of the objects, for example, but without limitation. For example, stacking relationships between boxes, overlap relationships between invoices, the state of a warped box, and relationships between shelves and boxes may be identified. The electronic device (101) may also identify relationships between invoice objects (1931, 1932, 1941) that are foreground objects, for example. For example, the overlap / separation state between invoice objects (1931, 1932, 1941) and / or the attachment state to box objects (1930, 1940) may be checked.For example, the electronic device (101) may reconstruct and provide a generated image based on physical arrangement relationships, and / or use it for performing tasks. For example, in the reconstructed image, invoice objects (1931, 1932) may be arranged so as to be spaced apart from each other. Or, if the boxes are distorted or irregularly stacked in the image, the boxes may be stacked in an orderly manner in the reconstructed image, but there are no limitations. The electronic device (101) may set at least one group, for example, using an LLM. For example, it may perform classification based on the analysis results of the invoice objects (1931, 1932, 1941) (which may be region or type of goods, but there are no limitations). The electronic device (101) may visually display the classification results or perform box movement tasks based on the classification results. The electronic device (101) may check additional information, for example (which may be item status information via access based on an administrator account, but there are no limitations). The electronic device (101) may, for example, provide text for boxes corresponding to categories based on region or type of goods, extract an invoice number, and / or extract the name of the goods written on the box. Those skilled in the art will understand that the inventory status of goods may be managed based on the information described above.
[0156] The electronic device (101) may include at least one processor (120) and a memory (130) for storing instructions.
[0157] When the above instructions are executed individually or collectively by the at least one processor (120), they may cause the electronic device (101) to provide a first image.
[0158] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to identify a first object on which a first group of text is written and a second object on which a second group of text is written, based on the confirmation of a request for text analysis for the first image.
[0159] When the above instructions are executed individually or collectively by the at least one processor (120), they may cause the electronic device (101) to verify the physical arrangement relationship in the real world between the first object and the second object.
[0160] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to provide an analysis result associated with a third text group including the first text group and the second text group, based on the determination that the first object and the second object together constitute an object in which a text group is written, based on the physical arrangement relationship.
[0161] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to provide an analysis result associated with either the first text group or the second text group based on the determination that the first object and the second object correspond to objects on which different text groups are written, based on the physical arrangement relationship.
[0162] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to determine that the first object and the second object correspond to objects on which different text groups are written, based on the fact that the layer corresponding to the first object and the layer corresponding to the second object are different.
[0163] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to determine that the first object and the second object constitute an object in which a text group is written, based on the fact that the shape of at least a part of the first object and the shape of at least a part of the second object correspond to each other.
[0164] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to determine that the first object and the second object constitute an object in which a text group is written, based on the fact that at least one attribute corresponding to the first object and at least one attribute corresponding to the second object are identical.
[0165] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to determine that the first object and the second object constitute an object in which a text group is written, based on the fact that the layer corresponding to the first object and the layer corresponding to the second object are identical.
[0166] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to provide a second image including a third object generated based on adjustments to at least some of the first object and the second object as an analysis result associated with the third text group, based on the determination that the first object and the second object together constitute an object in which a text group is written based on the physical arrangement relationship.
[0167] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to provide information identified based on a text recognition result for the third text group as at least part of an operation of providing an analysis result associated with the first text group and the second text group, based on the determination that the first object and the second object together constitute an object in which a text group is written, based on the physical arrangement relationship.
[0168] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to provide information related to text editing identified based on text recognition results and additional information for the third text group, as at least part of an operation of providing analysis results associated with the first text group and the second text group, based on the determination that the first object and the second object together constitute an object in which a text group is written based on the physical arrangement relationship.
[0169] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to select either the first object or the second object based on the analysis result of a third object different from the first object and the second object, as at least part of an operation of providing an analysis result associated with either the first text group or the second text group, based on the determination that the first object and the second object correspond to objects on which different text groups are written, based on the physical arrangement relationship.
[0170] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to select either the first object or the second object based on the analysis result of a third object different from the first object and the second object, and the recognition result for the first text group and / or the recognition result for the second text group, as at least part of an operation to provide an analysis result associated with either the first text group or the second text group, based on the determination that the first object and the second object each correspond to objects on which different text groups are written, based on the physical arrangement relationship.
[0171] The operation method of the electronic device (101) may include an operation that provides a first image.
[0172] The method of operation of the electronic device (101) may include, based on the confirmation of a request for text analysis for the first image, an operation of confirming a first object on which a first text group included in the first image is written and a second object on which a second text group is written.
[0173] The method of operation of the electronic device (101) may include an operation to verify the physical arrangement relationship in the real world between the first object and the second object.
[0174] The method of operation of the electronic device (101) may include providing an analysis result associated with a third text group including the first text group and the second text group based on the determination that the first object and the second object together constitute an object in which a single text group is written based on the physical arrangement relationship, or providing an analysis result associated with either the first text group or the second text group based on the determination that the first object and the second object each correspond to objects in which different text groups are written based on the physical arrangement relationship.
[0175] The method of operation of the electronic device (101) may further include an operation of determining that the first object and the second object correspond to objects on which different text groups are written, based on the fact that the layer corresponding to the first object and the layer corresponding to the second object are different.
[0176] The method of operation of the electronic device (101) may further include an operation of determining that the first object and the second object constitute an object on which a text group is written, based on the fact that the shape of at least a part of the first object and the shape of at least a part of the second object correspond to each other.
[0177] The method of operation of the electronic device (101) may further include an operation of determining that the first object and the second object constitute an object on which a text group is written, based on the fact that at least one attribute corresponding to the first object and at least one attribute corresponding to the second object are identical.
[0178] The method of operation of the electronic device (101) may further include an operation of determining that the first object and the second object constitute an object in which a single text group is written, based on the fact that the layer corresponding to the first object and the layer corresponding to the second object are identical.
[0179] Based on the determination that the first object and the second object together constitute an object in which a single text group is written based on the physical arrangement relationship above, the operation of providing an analysis result associated with a third text group including the first text group and the second text group may include providing a second image including a third object generated based on adjustment of at least some of the first object and the second object as an analysis result associated with the third text group.
[0180] Based on the determination that the first object and the second object correspond to objects on which different text groups are written, based on the physical arrangement relationship above, the operation of providing an analysis result associated with either the first text group or the second text group may include an operation of selecting either the first object or the second object based on the analysis result of a third object that is different from the first object and the second object.
[0181] Based on the determination that the first object and the second object correspond to objects on which different text groups are written, based on the physical arrangement relationship above, the operation of providing an analysis result associated with either the first text group or the second text group may include an operation of selecting either the first object or the second object based on the analysis result of a third object different from the first object and the second object, and the recognition result for the first text group and / or the recognition result for the second text group.
[0182] A storage medium for storing computer-readable instructions may be provided.
[0183] When the above instructions are executed individually or collectively by at least one processor (120) of the electronic device (101), the electronic device (101) may cause the electronic device (101) to provide a first image.
[0184] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to identify a first object on which a first group of text is written and a second object on which a second group of text is written, based on the confirmation of a request for text analysis for the first image.
[0185] When the above instructions are executed individually or collectively by the at least one processor (120), they may cause the electronic device (101) to verify the physical arrangement relationship in the real world between the first object and the second object.
[0186] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to provide an analysis result associated with a third text group including the first text group and the second text group, based on the determination that the first object and the second object together constitute an object in which a text group is written, based on the physical arrangement relationship.
[0187] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to provide an analysis result associated with either the first text group or the second text group based on the determination that the first object and the second object correspond to objects on which different text groups are written, based on the physical arrangement relationship.
[0188] The electronic device (101) may include at least one processor (120) and a memory (130) for storing instructions.
[0189] When the above instructions are executed individually or collectively by the at least one processor (120), they may cause the electronic device (101) to provide a first image.
[0190] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to identify a first object on which a first text group is written and a second object on which a second text group is written, based on the confirmation of a request for text analysis for the first image.
[0191] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to identify at least a portion of the first text group or the second text group as text to be recognized, based on the physical arrangement relationship between the first object and the second object.
[0192] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may cause to provide an analysis result associated with the recognition target text and a second image including a third object generated based on adjustments to at least some of the first object and the second object.
[0193] The method of operation of the electronic device (101) may include an operation of providing a first image.
[0194] The method of operation of the electronic device (101) may include the operation of verifying a first object on which a first text group is written and a second object on which a second text group is written, based on the verification of a request for text analysis for the first image.
[0195] The method of operation of the electronic device (101) may include an operation of identifying at least a portion of the first text group or the second text group as text to be recognized, based on the physical arrangement relationship between the first object and the second object.
[0196] The method of operation of the electronic device (101) may include providing an analysis result associated with a second image and a text to be recognized, the second image including a third object generated based on adjustment of at least some of the first object and the second object.
[0197] A storage medium for storing computer-readable instructions may be provided.
[0198] When the above instructions are executed individually or collectively by at least one processor (120) of the electronic device (101), the electronic device (101) may cause the electronic device (101) to provide a first image.
[0199] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to identify a first object on which a first text group is written and a second object on which a second text group is written, based on the confirmation of a request for text analysis for the first image.
[0200] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to identify at least a portion of the first text group or the second text group as text to be recognized, based on the physical arrangement relationship between the first object and the second object.
[0201] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may cause to provide an analysis result associated with the recognition target text and a second image including a third object generated based on adjustments to at least some of the first object and the second object.
[0202] The electronic device (101) may include at least one processor (120) and a memory (130) for storing instructions.
[0203] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may cause to provide a first image including a first object on which a first text group is written and a second object on which a second text group is written.
[0204] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to provide a second image including a third object created based on adjustments to at least some of the first object and the second object, and a third text group written on the third object, and an analysis result associated with the third text group, based on confirmation of a request for text analysis for the first image.
[0205] Here, the third text group may be generated based on at least some of the first text group or the second text group.
[0206] The method of operation of the electronic device (101) may include providing a first image including a first object on which a first text group is written and a second object on which a second text group is written.
[0207] The method of operation of the electronic device (101) may include, based on confirmation of a request for text analysis for the first image, providing a second image including a third object created based on adjustment of at least some of the first object and the second object, and a third text group written on the third object, and an analysis result associated with the third text group.
[0208] Here, the third text group may be generated based on at least some of the first text group or the second text group.
[0209] A storage medium for storing computer-readable instructions may be provided.
[0210] When the above instructions are executed individually or collectively by at least one processor (120) of the electronic device (101), the electronic device (101) may cause the electronic device (101) to provide a first image including a first object on which a first text group is written and a second object on which a second text group is written.
[0211] When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to provide a second image including a third object created based on adjustments to at least some of the first object and the second object, and a third text group written on the third object, and an analysis result associated with the third text group, based on confirmation of a request for text analysis for the first image.
[0212] Here, the third text group may be generated based on at least some of the first text group or the second text group.
[0213] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.
[0214] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0215] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0216] One embodiment of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0217] According to one embodiment, the method according to the embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0218] According to one embodiment, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to one embodiment, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to one embodiment, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
In the electronic device (101), At least one processor (120); and The electronic device (101) includes a memory (130) for storing instructions, wherein the instructions are executed individually or collectively by the at least one processor (120): Provide the first image, and Based on the confirmation of the request for text analysis regarding the first image above, the first object on which the first text group included in the first image is written and the second object on which the second text group is written are identified, and Identify the physical arrangement relationship in the real world between the first object and the second object, and Based on the determination that the first object and the second object together constitute an object in which a single text group is written, based on the physical arrangement relationship described above, an analysis result associated with a third text group including the first text group and the second text group is provided. An electronic device (101) that causes to provide an analysis result associated with either the first text group or the second text group, based on the determination that the first object and the second object correspond to objects on which different text groups are written, based on the physical arrangement relationship above. In Article 1, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) is: An electronic device (101) that causes the first object and the second object to be determined to correspond to objects on which different text groups are written, based on the fact that the layer corresponding to the first object and the layer corresponding to the second object are different. In any one of paragraphs 1 to 2, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) is: An electronic device (101) that causes the first object and the second object to be determined to constitute an object on which a text group is written, based on the fact that the shape of at least a part of the first object and the shape of at least a part of the second object correspond to each other. In any one of paragraphs 1 to 3, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) is: An electronic device (101) that causes the first object and the second object to be determined to constitute an object on which a text group is written, based on the fact that at least one attribute corresponding to the first object and at least one attribute corresponding to the second object are identical. In any one of paragraphs 1 to 4, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) is: An electronic device (101) that causes the first object and the second object to be determined to constitute an object on which a single text group is written, based on the fact that the layer corresponding to the first object and the layer corresponding to the second object are identical. In any one of paragraphs 1 to 5, The above instructions, when executed individually or collectively by the at least one processor (120), cause the electronic device (101) to provide an analysis result associated with a third text group including the first text group and the second text group based on the determination that the first object and the second object together constitute an object in which a text group is written based on the physical arrangement relationship, at least as part of an operation: An electronic device (101) that causes to provide a second image including a third object generated based on adjustments to at least some of the first object and the second object, along with an analysis result associated with the third text group. In any one of paragraphs 1 to 6, The above instructions, when executed individually or collectively by the at least one processor (120), cause the electronic device (101) to provide an analysis result associated with a third text group including the first text group and the second text group based on the determination that the first object and the second object together constitute an object in which a text group is written based on the physical arrangement relationship, at least as part of an operation: An electronic device (101) that causes to provide verified information based on the text recognition result for the third text group above. In any one of paragraphs 1 through 7, The above instructions, when executed individually or collectively by the at least one processor (120), cause the electronic device (101) to provide an analysis result associated with a third text group including the first text group and the second text group based on the determination that the first object and the second object together constitute an object in which a text group is written based on the physical arrangement relationship, at least as part of an operation: An electronic device (101) that causes to provide information related to text editing identified based on text recognition results and additional information for the above third text group. In any one of paragraphs 1 through 8, The above instructions, when executed individually or collectively by the at least one processor (120), cause the electronic device (101) to provide an analysis result associated with either the first text group or the second text group based on the determination that, based on the physical arrangement relationship, the first object and the second object each correspond to objects on which different text groups are written: An electronic device (101) that causes one of the first object and the second object to be selected based on the analysis results of a third object different from the first object and the second object. In any one of paragraphs 1 through 9, The above instructions, when executed individually or collectively by the at least one processor (120), cause the electronic device (101) to provide an analysis result associated with either the first text group or the second text group based on the determination that, based on the physical arrangement relationship, the first object and the second object each correspond to objects on which different text groups are written: An electronic device (101) that causes one of the first object and the second object to be selected based on the analysis result of a third object different from the first object and the second object, and the recognition result of the first text group and / or the recognition result of the second text group. In the method of operating the electronic device (101), Action of providing the first image; Based on the confirmation of a request for text analysis regarding the first image, an operation to identify a first object on which a first text group included in the first image is written and a second object on which a second text group is written; An operation to verify the physical arrangement relationship in the real world between the first object and the second object; An operation to provide an analysis result associated with a third text group including the first text group and the second text group, based on the determination that the first object and the second object together constitute an object on which a single text group is written, based on the physical arrangement relationship above, or to provide an analysis result associated with either the first text group or the second text group, based on the determination that the first object and the second object each correspond to objects on which different text groups are written, based on the physical arrangement relationship above. A method of operation of an electronic device (101) including In Article 11, An operation of determining that the first object and the second object each correspond to objects on which different text groups are written, based on the fact that the layer corresponding to the first object and the layer corresponding to the second object are different. A method of operation of an electronic device (101) including further In any one of paragraphs 11 to 12, An operation of determining that the first object and the second object constitute an object on which a single text group is written, based on the fact that the shape of at least a part of the first object and the shape of at least a part of the second object correspond to each other. A method of operation of an electronic device (101) including further In any one of paragraphs 11 to 13, An operation of determining that the first object and the second object constitute an object on which a single text group is written, based on the fact that at least one attribute corresponding to the first object and at least one attribute corresponding to the second object are identical. A method of operation of an electronic device (101) including further In any one of paragraphs 11 to 14, An operation of determining that the first object and the second object constitute an object on which a single text group is written, based on the fact that the layer corresponding to the first object and the layer corresponding to the second object are identical. A method of operation of an electronic device (101) including further