Method and system for enhancing image understanding

KR103024971B1Active Publication Date: 2026-09-29NAVER CLOUD CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
KR1020230117309
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-09-04
Publication Date
2026-09-29
Estimated Expiration
2043-09-04

Smart Images

  • Figure 112023097720633-PAT00024_ABST
    Figure 112023097720633-PAT00024_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for enhancing image understanding, performed by at least one processor. The method comprises the steps of receiving an image, receiving a text prompt, a first encoder generating a first set of embeddings based on the image, a second encoder generating a second set of embeddings based on information extracted from the image, and a decoder generating a first output based on the first set of embeddings, the second set of embeddings, and the text prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to a method and system for improving image understanding, and specifically, to a method and system in which a decoder generates a first output based on a first set of embeddings generated by a first encoder based on a received image, a second set of embeddings generated by a second encoder based on information extracted from the image, and a text prompt. Background Technology

[0002] Visually-situated Natural Language Understanding (NLU) combines computer vision and natural language processing to enable more precise analysis of visual data through language. Conventional technology has primarily focused on performing Optical Character Recognition (OCR) on visual document data and utilizing the extracted text for analysis. Furthermore, conventional technology has sought methods to process document images directly without relying on external OCR models. For example, conventional technology proposed a model that performs text reading directly from document images through pre-training, thereby enabling document understanding without an external OCR model.

[0003] Meanwhile, conventional technology has expanded training data and model parameters, and by adjusting a pre-trained Large Language Model (LLM) to match user intent, it can provide contextually appropriate answers to user queries in various tasks. Furthermore, conventional technology has made various attempts to integrate visual information into LLMs to process visual language tasks. For example, LLMs have been utilized to extract features from visual information via vision encoders and to process various tasks as inference modules.

[0004] While these conventional techniques have demonstrated successful vision language learning through LLM, they suffer from poor performance in NLU tasks placed in visual situations, such as question-and-answer interactions with visual documents. Furthermore, these conventional techniques have limitations in extracting detailed features from images. The problem to be solved

[0005] The present disclosure provides a method and system (device) for improving image understanding to solve the above-mentioned problems. means of solving the problem

[0006] The present disclosure may be implemented in various ways, including a method, an apparatus (system), or a computer program stored on a readable storage medium.

[0007] According to one embodiment of the present disclosure, an image understanding enhancement method may include the steps of receiving an image, receiving a text prompt, a first encoder generating a first set of embeddings based on the image, a second encoder generating a second set of embeddings based on information extracted from the image, and a decoder generating a first output based on the first set of embeddings, the second set of embeddings, and the text prompt.

[0008] A computer program stored on a computer-readable recording medium may be provided to execute a method according to one embodiment of the present disclosure on a computer.

[0009] A system according to one embodiment of the present disclosure comprises a communication module, a memory, and at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, wherein the at least one program may include instructions for receiving an image and receiving a text prompt, for a first encoder to generate a first set of embeddings based on the image, for a second encoder to generate a second set of embeddings based on information extracted from the image, and for a decoder to generate a first output based on the first set of embeddings, the second set of embeddings, and the text prompt. Effects of the invention

[0010] According to some embodiments of the present disclosure, the visual understanding of an image can be enhanced by using two encoders in the model. Additionally, by using Contrastive Learning (CL) to align embeddings encoded by two different encoders into a common feature space, the target referenced by the decoder is not biased, thereby improving the performance of the model. Accordingly, the model according to the present disclosure can derive answers to natural language questions more accurately from text-rich images.

[0011] According to some embodiments of the present disclosure, the integration of a neural network model and a language model can be made more flexible by incorporating a learned query mechanism and considering the application context. Accordingly, the language model can generate a more accurate and context-appropriate response while focusing on specific aspects of the visual input. This approach not only improves the language model's understanding of the visual context but also reduces computational costs as the learned query can efficiently extract relevant information from the input image.

[0012] According to some embodiments of the present disclosure, by utilizing an auxiliary encoder and a contrastive learning technique in a neural network model, the performance of a neural network model in a visual language understanding problem can be improved.

[0013] According to some embodiments of the present disclosure, when a contrastive learning technique is implemented, the embeddings of two encoders can be aligned more appropriately. Additionally, the embeddings of the two encoders can be distributed more randomly (or widely) within a common space. That is, the contrastive learning technique can significantly contribute to improving the performance of a neural network model.

[0014] The effects of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art to which the present disclosure pertains (referred to as "person skilled in the art") from the description in the claims. Brief explanation of the drawing

[0015] Embodiments of the present disclosure will be described with reference to the accompanying drawings described below, wherein similar reference numerals indicate similar elements, but are not limited thereto. FIG. 1 illustrates an example of a method for improving image understanding according to one embodiment of the present disclosure. FIG. 2 is a schematic diagram showing a configuration in which an information processing system is connected to communicate with a plurality of user terminals to improve image understanding according to one embodiment of the present disclosure. FIG. 3 is a block diagram showing the internal configuration of a user terminal and an information processing system according to one embodiment of the present disclosure. FIG. 4 is a diagram showing an overview of a standalone neural network model according to one embodiment of the present disclosure. FIG. 5 is a diagram showing an overview of a neural network model combined with a language model according to one embodiment of the present disclosure. FIG. 6 is a drawing showing an example of a contrasting feature alignment according to one embodiment of the present disclosure. FIG. 7 is a drawing showing an example of an embedding set of an auxiliary encoder according to one embodiment of the present disclosure. FIG. 8 is a diagram showing an example of a learning task of a model according to one embodiment of the present disclosure. FIG. 9 is a drawing showing an example of a graph illustrating the robustness of a model according to the OCR omission rate according to one embodiment of the present disclosure. FIG. 10 is a diagram showing the results of principal component analysis of the common space of embeddings generated by a vision encoder and an auxiliary encoder according to one embodiment of the present disclosure. FIG. 11 is a drawing showing an example of a graph illustrating the robustness of a neural network model according to one embodiment of the present disclosure. FIG. 12 is a diagram showing an example of a graph illustrating the cosine similarity of embeddings in a common space according to one embodiment of the present disclosure. FIG. 13 is a flowchart illustrating an example of a method according to one embodiment of the present disclosure. Specific details for implementing the invention

[0016] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the attached drawings. However, in the following description, specific descriptions regarding well-known functions or configurations will be omitted if there is a risk that the gist of the present disclosure may be unnecessarily obscured.

[0017] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Additionally, in the description of the following embodiments, the description of identical or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such component is not included in any embodiment.

[0018] The advantages and features of the disclosed embodiments and the methods for achieving them will become clear by referring to the embodiments described below in conjunction with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below but may be implemented in various different forms, and the embodiments provided are merely to make the present disclosure complete and to fully inform those skilled in the art of the scope of the invention.

[0019] The terms used in this specification will be briefly explained, and the disclosed embodiments will be described in detail. The terms used in this specification have been selected to be as generally used as possible, taking into account their functions in this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should be defined not merely by their names, but based on their meanings and the content throughout this disclosure.

[0020] In this specification, singular expressions include plural expressions unless the context clearly specifies them as singular. Additionally, plural expressions include singular expressions unless the context clearly specifies them as plural. Throughout the specification, when a part is described as including a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0021] Additionally, the terms 'module' or 'part' as used in the specification refer to software or hardware components, and the 'module' or 'part' performs certain roles. However, the meaning of 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside in an addressable storage medium or configured to run on one or more processors. Thus, as an example, the 'module' or 'part' may include components such as software components, object-oriented software components, class components, and task components, and at least one of processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, or variables. The components and the functions provided within the 'module' or 'part' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.

[0022] According to one embodiment of the present disclosure, a ‘module’ or ‘part’ may be implemented as a processor and memory. The term ‘processor’ should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, the term ‘processor’ may refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc. The term ‘processor’ may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors combined with a DSP core, or any other combination of such configurations. Additionally, the term ‘memory’ should be broadly interpreted to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as Random Access Memory (RAM), Read-Only Memory (ROM), Non-Volatile Random Access Memory (NVRAM), Programmable Read-Only Memory (PROM), Erasable-Programmable Read-Only Memory (EPROM), Electrically Erasable PROM (EEPROM), Flash Memory, Magnetic or Optical Data Storage Devices, Registers, etc. If a processor can read information from memory and / or write information to memory, the memory is said to be in an electronic communication state with the processor. Memory integrated into a processor is in an electronic communication state with the processor.

[0023] In the present disclosure, the 'system' may include at least one of a server device and a cloud device, but is not limited thereto. For example, the system may be composed of one or more server devices. As another example, the system may be composed of one or more cloud devices. As yet another example, the system may be configured and operated with both a server device and a cloud device.

[0024] In the present disclosure, 'display' may refer to any display device associated with a computing device, for example, any display device capable of displaying any information / data controlled by or provided by the computing device.

[0025] In the present disclosure, 'each of a plurality of A' or 'each of a plurality of A' may refer to each of all components included in a plurality of A, or each of some components included in a plurality of A.

[0026] In the present disclosure, a ‘machine learning model’ may include any model used to infer an answer to a given input. According to one embodiment, a machine learning model may include an artificial neural network model comprising an input layer, a plurality of hidden layers, and an output layer. Here, each layer may include a plurality of nodes. In the present disclosure, each of the plurality of machine learning models is described as a separate machine learning model, but is not limited thereto, and some or all of the plurality of machine learning models may be implemented as a single machine learning model. Additionally, a single machine learning model may include a plurality of machine learning models. In the present disclosure, the terms machine learning model and artificial neural network model may be used interchangeably to refer to the same or similar models.

[0027] In the present disclosure, a "Large Language Model (LLM)" may refer to a language model capable of inference without fine-tuning using methods such as few-shot learning, and may have more than 10 times as many parameters as a conventional general language model (e.g., more than 100 billion parameters). For example, a Large Language Model may be HyperCLOVA or GPT 3 (Generative Pretrained Transformer 3). In the present disclosure, a "language model" may include a Large Language Model.

[0028] With the advancement of Large Language Models (LLMs), research seeking to expand into the visual domain is surging. While conventional technologies demonstrate potential in generating concise captions for images and conducting natural language conversations, their performance regarding text-rich images is poor. This disclosure provides a Contrastive Reading Model (Cream), a novel neural structure designed to enhance the language-image understanding capabilities of LLMs by capturing complex details typically overlooked in conventional technologies. The method of this disclosure supports more effective understanding of text information within document images by integrating a vision encoder and an auxiliary encoder and supplementing them with contrastive feature alignment techniques. Consequently, the method of this disclosure bridges the gap between vision and language understanding, serving as a foundation for more sophisticated Document Intelligence Assistants. Furthermore, as a state-of-the-art model in the field of visual document understanding, the method of this disclosure provides superior performance for various tasks, such as Visual Question Answering (VQA) on document images.

[0029] Recent advancements in LLM have facilitated the development of numerous applications, providing valuable and meaningful services to users. Additionally, there is a growing number of studies extending these unimodal LLMs into multimodal LLMs, particularly large visual language models (LVLMs), by utilizing vision encoders to process information-rich visual tasks.

[0030] Various downstream tasks, such as image captioning, visual dialogue, evidence finding, reasoning, and question generation, can be used to evaluate LVLM. According to conventional technology, LVLM has shown limitations when processing text-rich visual tasks in these areas, resulting in poor applicability in real-world applications such as Document Visual Question Answering (DocVQA). Visual document understanding (VDU) tasks require the comprehensive analysis of various information types, including text, objects (e.g., graphs and charts), and layouts. However, because existing LVLMs can only extract granular features from images, it is difficult to provide a satisfactory solution in situations involving text-rich visual tasks.

[0031] The present disclosure provides a contrast reading model specifically designed to effectively overcome these limitations. The method of the present disclosure may include a simplified and practical architecture that fully integrates a vision encoder, an auxiliary encoder, and innovative learning techniques. In addition to a vision encoder for extracting overall visual features from document images, the method of the present disclosure may use auxiliary encoders, such as Optical Character Recognition (OCR) and object detectors, for extracting text and object-specific features. By utilizing the auxiliary encoder and the vision encoder, the method of the present disclosure can extract granular features without missing image details while understanding the visual context. Furthermore, the method of the present disclosure can be combined with LLM to overcome the limitations of LVLM and achieve superior performance in text-heavy visual tasks. Additionally, the present disclosure provides a contrast feature alignment method to mitigate bias between features extracted by each encoder during training to further enhance the performance of the model.

[0032] Experiments conducted on various VQA tasks for both the standalone model and the model combined with a frozen LLM show that the standalone model can achieve results comparable to the latest technology in tasks requiring the extraction of specific text information from document images. Additionally, when combined with an LLM, the model can demonstrate powerful performance in VDU tasks that are difficult for existing LVLMs to handle.

[0033] The present disclosure provides a learning technique associated with a novel model architecture tailored for visual document understanding tasks, which acts as the "eyes" of an LLM for performing text-rich tasks and can provide the LLM with both visual context and image details. By this configuration, the method of the present disclosure can achieve superior performance in various downstream tasks requiring the extraction of text information from document images. Furthermore, the method of the present disclosure can significantly improve the performance of specific downstream tasks by integrating with an LLM.

[0034] FIG. 1 illustrates an example of a method for improving image understanding according to one embodiment of the present disclosure. As illustrated, an image (110) and a text prompt (120) may be input to a model (130). Here, the text prompt (120) may be a natural language question associated with the image (110). The model (130) may generate an answer (140) to the natural language question based on the received image (110) and the text prompt (120).

[0035] In one embodiment, the model (130) may include a first encoder, a second encoder, and a decoder. Here, the first encoder may generate a first set of embeddings based on an image (110). Specifically, the first encoder may divide the image (110) into a plurality of patches. Additionally, the first encoder may generate a first set of embeddings based on the plurality of patches.

[0036] In one embodiment, the second encoder can generate a second set of embeddings based on information (specifically, text information) extracted from the image (110). Here, the information extracted from the image (110) may include first text information associated with an OCR box detected from the image (110) and second text information associated with an object box detected from the image (110). Examples of the first text information and the second text information are described in detail later with reference to FIG. 7.

[0037] In one embodiment, a first set of embeddings generated by a first encoder and a second set of embeddings generated by a second encoder can be aligned in a common feature space using a contrastive feature alignment technique. Specifically, the first encoder and the second encoder can be trained to align related data to similar locations in the common feature space using a contrastive feature alignment technique. Here, a specific box detected by a detector as an area containing text or an object in an image (110) and a specific patch of the image (110) containing text or an object can be determined as a positive pair of related data. Other relationships can be determined as negative pairs of unrelated data. That is, a specific patch in the image (110) existing at the same physical location as the detected specific box is considered to be related to each other, and thus the corresponding data pair can be aligned to similar locations in the common feature space.

[0038] In one embodiment, when the model (130) is used independently, the decoder can generate a first output based on a first set of embeddings, a second set of embeddings, and a text prompt (120). Additionally, the model (130) can generate a sequence of tokens associated with an answer to a natural language question referencing an image based on the first output. For example, the model (130) can output "DOG HOUSE" as an answer (140) to a text prompt (120) in the image (110) asking "What is the name of the store that likely sells beer in this ad?".

[0039] In one embodiment, the model (130) may be used in combination with a language model. In this case, the decoder may generate a first output based on a first set of embeddings, a second set of embeddings, a text prompt (120), and a learned query. Additionally, the model (130) may generate a soft visual prompt based on the first output. Furthermore, the language model may generate a second output based on the visual prompt and the text prompt. Here, the second output may be an answer to a natural language question.

[0040] With this configuration, the visual understanding of the image can be enhanced by the model (130) using two encoders. Additionally, by using a contrastive learning technique to align embeddings encoded by two different encoders into a common feature space, the target referenced by the decoder is not biased, thereby improving the performance of the model. Accordingly, the model can derive answers to natural language questions more accurately from text-rich images.

[0041] FIG. 2 is a schematic diagram showing a configuration in which an information processing system (230) is connected to communicate with a plurality of user terminals (210_1, 210_2, 210_3) to improve image understanding according to one embodiment of the present disclosure. As illustrated, the plurality of user terminals (210_1, 210_2, 210_3) may be connected to an information processing system (230) capable of providing a question-and-answer service for a visual document through a network (220). Here, the plurality of user terminals (210_1, 210_2, 210_3) may include a user terminal that receives a question-and-answer service for a visual document.

[0042] In one embodiment, the information processing system (230) may include one or more server devices and / or databases capable of storing, providing, and executing computer-executable programs (e.g., downloadable applications) and data associated with providing question-and-answer services for visual documents, or one or more distributed computing devices and / or distributed databases based on cloud computing services.

[0043] The question and answer service for a visual document provided by the information processing system (230) may be provided to the user through a question and answer service application for a visual document, a web browser, a web browser extension, etc. installed on each of the multiple user terminals (210_1, 210_2, 210_3). For example, the information processing system (230) may provide information or perform corresponding processing in response to a request for an answer to a question received from the user terminals (210_1, 210_2, 210_3) through a question and answer service application for a visual document.

[0044] Multiple user terminals (210_1, 210_2, 210_3) can communicate with an information processing system (230) through a network (220). The network (220) can be configured to enable communication between the multiple user terminals (210_1, 210_2, 210_3) and the information processing system (230). Depending on the installation environment, the network (220) may be configured as a wired network such as Ethernet, Power Line Communication, telephone line communication devices and RS-serial communication, a mobile communication network, a Wireless LAN (WLAN), Wi-Fi, Bluetooth and ZigBee, or a combination thereof. The communication method is not limited and may include not only communication methods utilizing communication networks that the network (220) may include (e.g., mobile communication network, wired internet, wireless internet, broadcasting network, satellite network, etc.) but also short-range wireless communication between user terminals (210_1, 210_2, 210_3).

[0045] In FIG. 2, a mobile phone terminal (210_1), a tablet terminal (210_2), and a PC terminal (210_3) are illustrated as examples of user terminals, but are not limited thereto. The user terminals (210_1, 210_2, 210_3) may be any computing device capable of wired and / or wireless communication and capable of installing and running a question-and-answer service application for visual documents or a web browser, etc. For example, user terminals may include an AI speaker, a smartphone, a mobile phone, a navigation system, a computer, a laptop, a digital broadcasting terminal, a PDA (Personal Digital Assistants), a PMP (Portable Multimedia Player), a tablet PC, a game console, a wearable device, an IoT (Internet of Things) device, a VR (Virtual Reality) device, an AR (Augmented Reality) device, a set-top box, etc. Additionally, FIG. 2 illustrates three user terminals (210_1, 210_2, 210_3) communicating with an information processing system (230) through a network (220), but is not limited thereto, and may be configured so that a different number of user terminals communicate with an information processing system (230) through a network (220).

[0046] In FIG. 2, a configuration in which a user's request is transmitted to an information processing system (230) through a user terminal (210_1, 210_2, 210_3) is illustrated as an example, but is not limited thereto. A user's request may be provided to the information processing system (230) through an input device associated with the information processing system (230) without passing through the user terminal (210_1, 210_2, 210_3), and the result of processing the user's request may be provided to the user through an output device (e.g., a display, etc.) associated with the information processing system (230).

[0047] FIG. 3 is a block diagram showing the internal configuration of a user terminal (210) and an information processing system (230) according to an embodiment of the present disclosure. The user terminal (210) may refer to any computing device capable of executing applications, web browsers, etc., and capable of wired / wireless communication, and may include, for example, the mobile phone terminal (210_1), tablet terminal (210_2), PC terminal (210_3) of FIG. 2. As illustrated, the user terminal (210) may include a memory (312), a processor (314), a communication module (316), and an input / output interface (318). Similarly, the information processing system (230) may include a memory (332), a processor (334), a communication module (336), and an input / output interface (338). As illustrated in FIG. 3, the user terminal (210) and the information processing system (230) may be configured to communicate information and / or data through the network (220) using their respective communication modules (316, 336). Additionally, the input / output device (320) may be configured to input information and / or data to the user terminal (210) or output information and / or data generated from the user terminal (210) through the input / output interface (318).

[0048] The memory (312, 332) may include any non-transient computer-readable recording medium. According to one embodiment, the memory (312, 332) may include a permanent mass storage device such as ROM (read-only memory), a disk drive, a solid-state drive (SSD), or flash memory. As another example, a permanent mass storage device such as ROM, an SSD, flash memory, or a disk drive may be included in the user terminal (210) or information processing system (230) as a separate permanent storage device distinct from the memory. Additionally, an operating system and at least one program code may be stored in the memory (312, 332).

[0049] These software components may be loaded from a computer-readable recording medium separate from memory (312, 332). This separate computer-readable recording medium may include a recording medium that can be directly connected to the user terminal (210) and the information processing system (230), for example, a computer-readable recording medium such as a floppy drive, disk, tape, DVD / CD-ROM drive, or memory card. As another example, the software components may be loaded into memory (312, 332) via a communication module (316, 336) rather than a computer-readable recording medium. For example, at least one program may be loaded into memory (312, 332) based on a computer program installed by files provided through a network (220) by developers or a file distribution system that distributes installation files for the application.

[0050] The processor (314, 334) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (314, 334) by memory (312, 332) or a communication module (316, 336). For example, the processor (314, 334) may be configured to execute instructions received according to program code stored in a recording device such as memory (312, 332).

[0051] The communication module (316, 336) may provide a configuration or function for the user terminal (210) and the information processing system (230) to communicate with each other via the network (220), and may provide a configuration or function for the user terminal (210) and / or the information processing system (230) to communicate with another user terminal or another system (e.g., a separate cloud system). For example, a request or data (e.g., a request for an answer to a question) generated by the processor (314) of the user terminal (210) according to program code stored in a recording device such as memory (312) may be transmitted to the information processing system (230) via the network (220) under the control of the communication module (316). Conversely, a control signal or command provided under the control of the processor (334) of the information processing system (230) can be received by the user terminal (210) through the communication module (336) and the network (220) via the communication module (316) of the user terminal (210).

[0052] The input / output interface (318) may be a means for interfacing with an input / output device (320). As an example, the input device may include a device such as a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, or a mouse, and the output device may include a device such as a display, a speaker, or a haptic feedback device. As another example, the input / output interface (318) may be a means for interfacing with a device in which the configuration or function for performing input and output is integrated into one, such as a touchscreen. For example, when the processor (314) of the user terminal (210) processes instructions of a computer program loaded in memory (312), a service screen configured using information and / or data provided by an information processing system (230) or another user terminal may be displayed on a display through the input / output interface (318). In FIG. 3, the input / output device (320) is depicted as not being included in the user terminal (210), but is not limited thereto and may be configured as a single device with the user terminal (210). Additionally, the input / output interface (338) of the information processing system (230) may be a means for interfacing with a device (not shown) for input or output that is connected to the information processing system (230) or that the information processing system (230) may include. In FIG. 3, the input / output interface (318, 338) is shown as an element configured separately from the processor (314, 334), but is not limited thereto, and the input / output interface (318, 338) may be configured to be included in the processor (314, 334).

[0053] The user terminal (210) and the information processing system (230) may include more components than those of FIG. 3. However, it is not necessary to clearly illustrate most of the prior art components. In one embodiment, the user terminal (210) may be implemented to include at least some of the input / output devices (320) described above. Additionally, the user terminal (210) may further include other components such as a transceiver, a GPS (Global Positioning System) module, a camera, various sensors, a database, etc. For example, if the user terminal (210) is a smartphone, it may include components that are generally included in a smartphone, and may be implemented to include various components such as an accelerometer, a gyroscope, a microphone module, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration.

[0054] While a program or application for a question-and-answer service for a visual document is in operation, the processor (314) may receive text, images, video, voice and / or actions, etc. that are input or selected through an input device such as a touch screen, keyboard, audio sensor and / or image sensor, camera, microphone, etc. connected to an input / output interface (318), and may store the received text, images, video, voice and / or actions, etc. in memory (312) or provide them to an information processing system (230) through a communication module (316) and a network (220).

[0055] The processor (314) of the user terminal (210) may be configured to manage, process, and / or store information and / or data received from an input / output device (320), another user terminal, an information processing system (230), and / or a plurality of external systems. The information and / or data processed by the processor (314) may be provided to the information processing system (230) through a communication module (316) and a network (220). The processor (314) of the user terminal (210) may transmit information and / or data to the input / output device (320) through an input / output interface (318) to output it. For example, the processor (314) may output or display the received information and / or data on a screen associated with the user terminal (210).

[0056] The processor (334) of the information processing system (230) may be configured to manage, process, and / or store information and / or data received from a plurality of user terminals (210) and / or a plurality of external systems. The information and / or data processed by the processor (334) may be provided to the user terminals (210) through a communication module (336) and a network (220).

[0057] FIG. 4 is a diagram showing an overview of a standalone neural network model (420) according to one embodiment of the present disclosure. In one embodiment, the neural network model (420) may include a vision encoder (422), an auxiliary encoder (424), and a decoder (426). As illustrated, when the neural network model (420) is used independently, the decoder (426) can directly generate the requested information in the form of text.

[0058] Based on specific evidence within an image, it is necessary to derive accurate answers to natural language questions from the input image. For example, if a question requiring the extraction of specific information from a document image is input, even if the answer is linguistically plausible, it may be meaningless if the text within the image is not accurately recognized and processed. Therefore, it is necessary to accurately respond to a given natural language question by effectively utilizing the information contained in the image. In this process, it is important to identify specific factual evidence within the image, such as text, objects, and other relevant features.

[0059] In one embodiment, when an image (410) is input, the vision encoder (422) can generate a first set of embeddings based on the image (410). Additionally, specific feature evidence (e.g., text or objects within the image) can be extracted from the image (410) through a detector (e.g., OCR or object detector). In this case, the auxiliary encoder (424) can embed the extracted items into a common feature space. Here, the outputs of the vision encoder (422) and the auxiliary encoder (424) can be trained to align to the common feature space through a contrastive learning scheme. Then, the features can be concatenated in a sequence direction rather than a dimension direction and provided to the cross-attention layer of the decoder (426). Accordingly, the number of embeddings that the decoder (426) can reference can be doubled. Additionally, the decoder (426) can generate a first output through an attention mechanism based on the natural language question (430) and the embeddings output from the two encoders (422, 424), and based on the first output, an answer (440) associated with the natural language question (430) can be generated.

[0060] In one embodiment, the vision encoder (422) is an image (x ∈ RHХWХC )(410) is the first set of embeddings {z i |z i ∈ R d It can be encoded as {1 ≤ i ≤ n}. Here, n is the size of the feature map or the number of image patches (412), and d is the dimension of the resulting output vector. As the encoder network, a CNN (Convolutional Neural Network) based model or a transformer based model may be used. In this disclosure, for simplification, a vision transformer equipped with 2D absolute position encoding and a variable resolution mechanism was used, but is not limited thereto. Here, the variable resolution mechanism may refer to an input image preprocessing strategy that converts an image into a fixed number of patches without distorting the original image aspect ratio. For example, an input image (410) may be divided into a plurality of image patches (412) having the same aspect ratio as the image (410). That is, the vision encoder (422) may generate a first set of embeddings based on the plurality of image patches (412).

[0061] In one embodiment, the auxiliary encoder (424) contains information of feature evidence extracted from the image (410), such as the OCR box (416) and the universal object box (414), in a second set of embeddings { | ∈ R d , 1≤i≤ It can be encoded as follows. That is, the extracted feature evidence can be converted into a series of token embeddings. In this conversion, the recognized text can be used in the OCR box (416), and the recognized semantic object label can be used in the universal object box (414). Additionally, type embeddings may be added to distinguish between the OCR box (416) and the universal object box (414). Furthermore, 2D absolute position encoding (e.g., the center X, Y coordinates of the box, etc.) may be applied to encode position information. Here, the backbone may use a BART (Bidirectional and Auto-Regressive Transformers) encoder architecture, but is not limited thereto.

[0062] In one embodiment, the decoder (426) is formed from the embeddings {z1, ..., z generated from two encoders}z1, ..., z n , , ..., Process} and vector sequence h ∈ R mХd (i.e., the first output) can be generated. Here, BART can be used as the decoder architecture, and m can represent the sequence length of the generated vector. This vector sequence, referred to as the last hidden state of the decoder (426), can be associated with the answer (440) to the natural language question (430). Specifically, in the standalone neural network model (420), the weight matrix W ∈ R dХv A linear language modeling head represented by is applied to a hidden state, thereby forming a token sequence associated with the answer (440) to a natural language question (430). = hW can be generated. Here, ∈ R mХvis a predicted token sequence, and v may represent the size of the token vocabulary. Additionally, the decoder (426) may use an attention mechanism having a unidirectional attention flow. Furthermore, the neural network model (420) may adopt a language modeling loss so that the decoder (426) generates a conditional token embedding sequence from an image.

[0063] FIG. 5 is a diagram showing an overview of a neural network model (520) combined with a language model (560) according to one embodiment of the present disclosure. In one embodiment, when the neural network model (520) is combined with a language model (e.g., LLM) (560), the output of the decoder (526) can serve as a soft visual prompt (550). That is, a vector sequence of hidden states output by the decoder (526) can be used as a soft visual prompt (550) of the language model (560). The description of the configuration shown in FIG. 5 that was described above in FIG. 4 is omitted.

[0064] Language models (or LLMs) possess state-of-the-art performance in a wide range of natural language processing tasks, such as text classification, question answering, and machine translation. However, these language models have limitations in understanding and responding to contextual language. To address this problem, a neural network model (520) and a language model (560) can be integrated. In this case, features extracted by the decoder (526) of the neural network model (520) are used as a soft visual prompt (550) for the language model (560), and the language model (560) can generate an answer (570) to a given input image (510) and a natural language question (540).

[0065] In one embodiment, when the neural network model (520) is combined with the language model (560), the last hidden state (first output) of the decoder (526) is a weight matrix U ∈ R dХd' It can be linearly transformed (h' = hU) using , where d' can represent the dimension of the input embeddings of the language model. Then, the transformed hidden state h' ∈ R mХd' It can be used as a soft visual prompt (550) as an input to the language model (560) to combine the visual understanding ability of the neural network model (520) with the language processing ability of the language model (560). Here, the neural network model (520) may adopt a language modeling loss so that the decoder (526) generates a conditional token embedding sequence from an image.

[0066] In one embodiment, a learned query mechanism may be applied to enhance the integration of the neural network model (520) and the language model (560). This mechanism can extract a hidden state of a fixed size and provide it to the language model (560) by utilizing the learned query (530) as a set of learnable embeddings as the input to the decoder (526). Here, the number of vectors included in the learned query (530) may be equal to the number of vectors included in the soft visual prompt (550). That is, the learned query (530) may serve to fix the number of vectors in the soft visual prompt (550) to a predetermined number.

[0067] In one embodiment, a natural language question (540) is input to a decoder (526) along with a learned query (530), so that the last hidden state of the decoder (526) for the learned query (530) can encode more valuable information to answer the natural language question (540). In this way, the combination of the neural network model (520) and the language model (560) can perform a wider variety of roles in actual applications.

[0068] Through this configuration, the length of the soft visual prompt is shorter than when inputting all OCR tokens into the neural network model, which can reduce computational costs. Specifically, the complexity per attention layer is mathematically O( 2 It can be expressed as ), here represents the sequence length of the tokens, and represents the hidden dimension of the model. In particular, For LLMs with large values ​​and many attention points (e.g., 175B), reducing the input token length reduces complexity to a minuscule degree. Through this efficiency, neural network models can provide superior performance while consuming fewer resources.

[0069] Directly inputting OCR into an LLM can incur high computational costs. However, when a neural network model is integrated with an LLM, the length of visual prompts is reduced, allowing the neural network model to achieve better performance while using fewer tokens and computing resources. This efficiency highlights the potential of the neural network model of the present disclosure in visual language understanding tasks, particularly when compared to other LLM integration approaches.

[0070] With this configuration, the integration of neural network models and language models can become more flexible by incorporating learned query mechanisms and considering the application context. Consequently, the language model can generate more accurate and context-appropriate responses by focusing on specific aspects of the visual input. This approach not only improves the language model's understanding of the visual context but also reduces computational costs as the learned queries can efficiently extract relevant information from input images.

[0071] FIG. 6 is a diagram illustrating an example of contrastive feature alignment according to one embodiment of the present disclosure. In one embodiment, a vision encoder (640) may generate a first set of embeddings (642) based on an image (610). Specifically, the image (610) may be divided into a plurality of patches (612), and the vision encoder (640) may encode each of the plurality of patches (612) to generate a first set of embeddings (642). Additionally, an auxiliary encoder (650) may generate a second set of embeddings (652) based on the image (610). In this case, the embeddings (642, 652) generated by the two encoders (640, 650) may be aligned to a common feature space (660) using a contrastive learning technique.

[0072] In order to integrate information such as text and object data along with information of an image (610) within the decoder, text information can be encoded using an auxiliary encoder (650). However, it is uncertain whether features generated by different encoders (640, 650) will be well aligned in a common space. To address this, a contrast learning technique can be used in parallel with the training of the neural network model.

[0073] In one embodiment, given an OCR box and / or universal object box obtained from an OCR (630) and / or object detector (620), the embedding of the patch where the corresponding evidence (text or object) is physically located among a plurality of patches (612) of the image (610) and the embedding of the feature evidence may contain semantically similar information. Accordingly, by applying a contrastive learning technique, the relationship between the embedding of the feature evidence and the image patch corresponding to the evidence can be defined as a positive pair, and all other relationships can be defined as negative pairs.

[0074] For example, if an image contains a 'book' with the title 'Apple', the patch in the image where the book is located can form a positive pair with bounding box information labeled 'Book'. Additionally, an image containing the word 'Apple' can form a positive pair with the 'Apple' text label and its corresponding bounding box information. In this case, all other relationships can be formed as negative pairs. Through this contrastive learning technique, more relationship pairs can be obtained from a single sample compared to image-level contrastive learning approaches (e.g., CLIP).

[0075] In one embodiment, for the contrastive learning technique, a 2-layer multi-layer perceptron (MLP) : R d R d* This can be used. Here, d* can represent a hyperparameter for the dimension of the common space. The goal of the contrastive learning technique can be expressed as Equation 1 below.

[0076]

[0077] Here, {v i |1≤i≤l} and { Each of the |1≤i≤l} sets is z and It represents the stacked features in the order combined as positive pairs. Additionally, l represents the number of feature evidences, and is a temperature parameter that controls softmax sharpness. And, the function s(x, y) = cos( (x), (y)) is Cosine similarity between input vectors can be calculated using etherized MLP parameters. This contrastive learning technique can encourage the embeddings of two encoders to align in a common feature space.

[0078] FIG. 7 is a diagram illustrating an example of an embedding set of an auxiliary encoder (720) according to an embodiment of the present disclosure. In one embodiment, the auxiliary encoder (720) can generate embeddings based on features extracted from an input image (710), such as an OCR box and a universal object box. Specifically, the OCR can detect an area containing text in the image (710) as an OCR box. At this time, the OCR can extract / store location information of the OCR box containing text and text information recognized within the OCR box. Additionally, an object detector can detect an area containing a specific object in the image (710) as an object box. Here, the object detector can extract / store location information of the object box and detected object labels (e.g., Wear, Person, etc.). That is, the information extracted by the OCR and the object detector may all be text information.

[0079] In one embodiment, the extracted features may be converted into a series of token embeddings. Here, the generated embeddings may include type embeddings (740) for distinguishing between OCR boxes and universal object boxes, location information (730) indicating the 2D absolute position of the corresponding box in the input image (710), and text information / object labels (750) indicating the recognized features as sub-words. Here, the location information (730) may be the center coordinates (X, Y coordinates) of the corresponding box in the image.

[0080] FIG. 8 is a diagram illustrating an example of a learning task of a model according to one embodiment of the present disclosure. A first image (810) shows an example of an image used for learning. Additionally, a second image (820) shows an example of a response for each task generated by a neural network model during learning.

[0081] In one embodiment, a neural network model may perform a Text Read (TR) task for modeling text within an image. In this case, due to the introduction of an auxiliary encoder, some OCR tokens may be replaced with mask tokens when the OCR result is input into the auxiliary encoder. Accordingly, the neural network model may be trained to read not only the mask tokens but also the entire text of the image by simultaneously using image modalities.

[0082] In one embodiment, the neural network model may perform a Masked Text Prediction (MTP) task to predict hidden text in an image in order to improve the understanding of the overall context of the image. Accordingly, when some OCR boxes of the image are randomly masked, the neural network model may be trained to predict characters in areas that are completely obscured in the image.

[0083] In one embodiment, a neural network model may perform an image caption (Capt.) task to generate a natural language description of an image. This natural language description, i.e., the caption, comprehensively represents the image context and is very important in the task of language understanding regarding visual situations. By capturing the entire scene and object details of the image, the neural network model can be trained to understand the overall context of the image and recognize objects.

[0084] In one embodiment, a neural network model may perform a Question Answering (QA) task to generate appropriate answers to images and natural language questions. Through this task, the neural network model can learn how to provide more accurate answers by focusing on specific image regions and text information. Accordingly, the neural network model's understanding of the relationship between the visual information of an image and text information can be improved.

[0085] In one embodiment, a neural network model may perform a Question Generation (QG) task to generate a question corresponding to a given answer in the context of an image. Through this task, the neural network model's ability to answer questions may be enhanced by improving the ability to infer the content of the image and the answer. The question generation task may be performed simply by swapping the question and the answer.

[0086] In one embodiment, since the tasks described above, such as text reading, masked text prediction, image captioning, question answering, and question generation, are related to each other, a neural network model can solve the tasks using a similar approach. These tasks may include extracting a sequence of text based on a given task command query when an input image and feature evidence within the image are provided. For example, a query for text reading could be "Please read all the text from the top left to the bottom right of the image," and for masked text prediction, it could be "Please guess all the text hidden in the masked area." Additionally, for image captioning, queries such as "Please describe the image" or "Please explain the image" may be used.

[0087] The second image (820) shows an example of an integrated learning framework in which desired answer text is generated for all tasks when a natural language prompt (or query) and an image are input. Unlike existing document understanding methods that use prompts assigned to a single task, the neural network model is trained with natural language-based prompts and can be more seamlessly integrated into the language model.

[0088] In one embodiment, a neural network model can be trained by combining a supervised fine-tuning VQA dataset with pre-trained datasets for text reading, mask text prediction, and image captioning. Here, some question-answering data may require more inference than simply reading text in an image or describing the situation. Therefore, as shown in Table 1 below, the neural network model can be trained by limiting the use of these fine-tuning QA benchmarks during the initial training stage and increasing the weight of QA data after the intermediate stage.

[0089] Phase Task Proportion Cream-phase1 Text Read (22%), MTP (46%), Captioning (22%), QA (5%), QG (5%) Cream-phase2 Text Read (7%), MTP (14%), Captioning (26%), QA (48%), QG (5%) LLM Integration QA (100%)

[0090] In one embodiment, the training of the neural network model may consist of two main objectives: language modeling loss and contrast learning loss. These objectives may involve aligning the sets of embeddings generated by the vision encoder and the auxiliary encoder, and demonstrating the overall performance of the model in a language comprehension task in a visual context. Specifically, the objective of language modeling is to generate a sequence of token embeddings that corresponds appropriately to the image. By using a simple cross-entropy loss in the neural network model, the difference between the predicted token sequence and the actual data can be measured. Here, a teacher-forcing scheme may be utilized in the training process, which allows the model to learn through accurate contextual information by using the ground truth data as input instead of the model output from the previous time step.

[0091] In one embodiment, the goal of the contrastive learning technique may be to encourage alignment of embedding sets generated by a vision encoder and an auxiliary encoder in a common feature space. Such alignment is important for effectively integrating OCR information and object information along with image information within the decoder. To this end, positive and negative pairs are defined based on location information of feature evidence within the image. Positive pairs consist of relationships in which the feature evidence and the corresponding image patch share semantically similar information, while negative pairs may consist of any other relationships. The goal of the contrastive learning technique can be expressed as Equation 1 described above.

[0092] In one embodiment, to combine two goals during the training process of a neural network model, a language modeling loss (L) as shown in Equation 2 below LM ) and contrasting learning loss (L CL The weighted sum of ) can be used.

[0093]

[0094] Here, represents a hyperparameter that controls the relative importance of the two goals. By integrating these two learning goals, the neural network model can effectively align information encoded by two encoders and achieve high performance in the task of language understanding in visual situations.

[0095] In one embodiment, the decoder of the neural network model can function as a soft visual prompter. The transformed hidden state can serve as a soft visual prompt that adjusts the language model according to the visual representation extracted by the decoder of the neural network model. Throughout the integration process, both the encoders of the language model and the neural network model are frozen, and only the decoder of the neural network model can be updated through gradient descent-based training. In this case, a new learnable parameter called a vision query is introduced to extract a fixed number of vectors that serve as soft visual prompts. Given a desired number of vectors k, new k token embeddings that serve as vision queries can be generated. These queries are input into the decoder of the neural network model, and the resulting output vector can be used as a visual prompt.

[0096] In one embodiment, when integrated with a language model, the decoder of the neural network model may no longer function as an autoregressive decoder. To address this issue, the attention mechanism can be tuned to allow a bidirectional attention flow in the decoder. This tuning can provide a more efficient model for language comprehension tasks in visual contexts by enhancing the decoder's ability to effectively combine the neural network model's understanding of visual information with the language model's language processing capabilities.

[0097] FIG. 9 is a diagram showing an example of a graph (900) illustrating the robustness of a model according to an OCR omission rate according to an embodiment of the present disclosure. The graph (900) shows the robustness of the model according to an OCR omission rate that gradually increases for DocVQA samples. As can be seen in the graph (900), the neural network model (Cream) according to the present disclosure can provide robust performance even in situations where auxiliary information is not provided.

[0098] In graph (900) and Table 2 below, the neural network model of the present disclosure is "Cream" and "Cream small Neural network models marked with " and without contrastive learning techniques used are "Cream small Neural network models that are marked "w / o CL" and do not use an auxiliary encoder in the test are marked "disable aux. at test". Table 2 below shows the effect of auxiliary encoders and contrastive learning techniques on the performance of neural network models.

[0099] Model DocVQA ChartQA InfoVQA Cream Small 67.8 54.8 29.9 - disable aux. at test 38.9 41.4 13.1 Diff. 28.9 13.4 16.8 Cream Small w / o CL 65.3 52.7 31.5 - disable aux. at test 7.9 9.8 12.1 Diff. 57.4 42.9 19.4 Donut-like 49.8 47.6 16.8 Donut-like-Patch1700 60.5 52.2 20.9

[0100] The performance difference between the neural network model of the present disclosure and a vision-only model (e.g., a Donut-like model) indicates the importance of using an auxiliary encoder that integrates OCR and object detection results. This result can support the effectiveness of the approach according to the present disclosure, which integrates auxiliary information to enhance visual understanding.

[0101] The effectiveness of the proposed contrastive learning technique can be verified by comparing a model using the technique with one that does not. In the absence of contrastive learning, the vision encoder may be improperly trained during the inference phase, leading to performance degradation. The relatively small performance difference observed with contrastive learning demonstrates its ability to prevent feature collapse or bias, which highlights the importance of contrastive learning in the overall design of neural network models.

[0102] The gap between the difference between a neural network model and a neural network model without an auxiliary encoder, and between a neural network model without a contrast learning technique and a neural network model without an auxiliary encoder, indicates that the contrast learning technique contributes to balancing the two encoders, thereby providing more robust performance even in the absence of auxiliary information. This feature can have a very significant impact on actual applications. For example, in graph (900), as the OCR omission rate increases, it can be seen that the difference between a neural network model without an auxiliary encoder and a neural network model without a contrast learning technique increases.

[0103] The Donut-like-Patch1700 model is trained by increasing the number of patches from 1024 to 1700 in the Donut-like model to accommodate larger images. However, while this modification requires significant computational resources, it still results in lower scores than the neural network model according to the present disclosure. These results imply that it is difficult to bridge the performance gap between models with and without auxiliary encoding simply by scaling the vision encoder. By utilizing auxiliary encoders and contrast learning techniques in the neural network model with this configuration, the performance of the neural network model in visual language understanding problems can be improved.

[0104] FIG. 10 is a diagram showing the results of principal component analysis on the common space of embeddings generated by a vision encoder and an auxiliary encoder according to one embodiment of the present disclosure. A first image (1010) shows the results of principal component analysis (PCA) on the common space of embeddings generated by the two encoders when no contrastive learning technique is used. Additionally, a second image (1020) shows the results of principal component analysis on the common space of embeddings generated by the two encoders when a contrastive learning technique is used.

[0105] The PCA results for the space using the contrastive learning technique show improved alignment, particularly in the first principal component (PC1), which reflects the difference between the two encoders. When visualizing with the second principal component (PC2) and third principal component (PC3) excluding the first principal component, it is confirmed that the alignment of each embedding improves in the common space of the neural network model with the contrastive learning technique. This may indicate that similar embeddings are clustered more effectively.

[0106] Therefore, implementing contrastive learning techniques allows the embeddings of the two encoders to be aligned more appropriately. In other words, contrastive learning techniques can significantly contribute to improving the performance of neural network models.

[0107] FIG. 11 is a diagram showing an example of a graph illustrating the robustness of a neural network model according to one embodiment of the present disclosure. A first graph (1110) shows the difference in robustness between a neural network model according to the present disclosure (e.g., "Cream") and a conventional model (e.g., "UDOP") in various OCR engines (e.g., "Clova", "EASY", "Paddle"). According to the first graph (1110), it is confirmed that the neural network model has excellent resilience and can maintain performance despite performance differences among various OCR engines.

[0108] The second graph (1120) shows the effect of contrastive learning techniques across various OCR engines. According to the second graph (1120), it is confirmed that training using contrastive learning techniques can improve the robustness of neural network models and mitigate performance degradation despite performance differences among various OCR engines.

[0109] FIG. 12 is a diagram showing an example of a graph illustrating the cosine similarity of embeddings in a common space according to one embodiment of the present disclosure. The first graph (1210) shows an example in which two embeddings are randomly selected in the common feature space of a neural network model using a contrastive learning technique, their cosine similarity is calculated, and then a histogram is constructed. Additionally, the second graph (1220) shows an example in which two embeddings are randomly selected in the common feature space of a neural network model not using a contrastive learning technique, their cosine similarity is calculated, and then a histogram is constructed.

[0110] When dealing with random (unit) vectors in the embedding space, a Gaussian distribution can be expected in the histogram. According to the first graph (1210), when a contrastive learning technique is applied to the embedding space, the distribution of the embeddings appears wider, which indicates that the quality of the embedding space can be improved. In graphs (1210, 1220), the solid line represents a Gaussian distribution with minimum KL divergence for each histogram distribution. Specifically, each distribution is N(0, 1 / 140) when the contrastive learning technique is applied, and N(0, 1 / 3) when the contrastive learning technique is not applied.

[0111] Therefore, the embeddings of the two encoders can be distributed more randomly (or widely) within a common space. In other words, contrastive learning techniques can significantly contribute to improving the performance of neural network models.

[0112] FIG. 13 is a flowchart illustrating an example of a method (1300) according to one embodiment of the present disclosure. In one embodiment, the method (1300) may be performed by at least one processor. The method (1300) may be initiated by the processor receiving an image (S1310). Additionally, the processor may receive a text prompt (S1320).

[0113] In one embodiment, a first encoder can generate a first set of embeddings based on an image (S1330). Specifically, the processor can divide the image into a plurality of patches. Here, the aspect ratio of each of the plurality of patches may be the same as the aspect ratio of the image. Additionally, the first encoder can generate a first set of embeddings based on the plurality of patches.

[0114] In one embodiment, a second encoder may generate a second set of embeddings based on information extracted from an image (S1340). Here, the second set of embeddings may include a first embedding and a second embedding. Specifically, a first detector may detect an area containing text in an image as an OCR box. In this case, the second encoder may generate a first embedding based on first text information associated with the OCR box detected from the image. Here, the first text information associated with the OCR box may include location information of the OCR box, recognized text information, and text information representing the first detector. Additionally, the second detector may detect an area containing a specific object in an image as an object box. In this case, the second encoder may generate a second embedding based on second text information associated with the object box detected from the image. Here, the second text information associated with the object box may include location information of the object box, recognized object label, and text information representing the second detector.

[0115] In one embodiment, the decoder can generate a first output based on a first set of embeddings, a second set of embeddings, and a text prompt (S1350). Specifically, the first set of embeddings generated by the first encoder and the second set of embeddings generated by the second encoder can be concatenated in a sequence direction and input to the cross-attention layer of the decoder.

[0116] In one embodiment, the first encoder and the second encoder may be trained to align related data to similar locations in a common feature space using a contrastive feature alignment technique. Here, a specific box detected by the detector as an area containing text or an object in an image and a specific patch of the image containing the text or object may be determined as a related data pair.

[0117] In one embodiment, the processor may divide an image into a plurality of patches. Here, the plurality of patches may include a first patch containing an OCR box and a second patch containing an object box. Additionally, a first encoder may generate a third embedding associated with the first patch. Furthermore, the first encoder may generate a fourth embedding associated with the second patch. In this case, the first embedding and the third embedding may be aligned to a first position in the common feature space, and the second embedding and the fourth embedding may be aligned to a second position in the common feature space.

[0118] In one embodiment, the text prompt may be a natural language question associated with an image. In this case, the processor may generate a token sequence associated with an answer to the natural language question referencing the image based on the first output. Here, the decoder may use an attention mechanism having a unidirectional attention flow.

[0119] In one embodiment, the decoder may generate a first output based on a first set of embeddings, a second set of embeddings, a text prompt, and a learned query. In this case, the processor may generate a soft visual prompt based on the first output. Additionally, the language model may generate a second output based on the soft visual prompt and the text prompt. Here, the number of vectors included in the learned query and the number of vectors included in the soft visual prompt may be the same. Furthermore, the decoder may use an attention mechanism having a bidirectional attention flow. And, while the first encoder, the second encoder, and the language model are in a frozen state, the decoder may be updated with additional learning.

[0120] The present disclosure provides a neural network model designed to address the limitations of existing LVLMs in text-heavy visual tasks. This neural network model features a simplified and practical architecture that seamlessly integrates a vision encoder and an auxiliary encoder, as well as innovative learning techniques including a contrastive feature alignment method. Through extensive testing on language comprehension tasks in various visual contexts, the neural network model of the present disclosure provides state-of-the-art performance in tasks requiring the extraction of text information from document images. Table 3 below shows the results of various models performing language comprehension tasks in visual contexts on various benchmarks (e.g., DocVQA, ChartQA, Info VQA).

[0121] Model Prompt Length Use Auxiliary DocVQA ChartQA InfoVQA OCR-Vicuna7B |OCR| o 29.2 6.2 13.6 OCR-Vicuna13B |OCR| o 31.4 3.7 23.7 OCR-GPT3.5 |OCR| o 62.4 15.9 26.6 OCR-GPT4 |OCR| o 75.9 34.3 25.0 BLIP2-OPT-6.7B 32 - 3.7 4.6 11.0 BLIP2xOCR-OPT-6.7B 32+|OCR| o 6.2 17.5 30.4 BLIP2-FlanT5xxL-11B 32 - 8.6 4.4 11.4 BLIP2xOCR-FlanT5xxL-11B 32+|OCR| o 63.8 18.3 36.6 LLaVA-Vicuna7B 256 5.5 0.5 2.4 LLaVA-Vicuna13B 256 5.9 1.4 3.1 Cream-Vicuna7B ( Proposed ) 192 o 80.0 61.6 42.4

[0122] In benchmarks requiring enhanced visual understanding capabilities, the integration of a neural network model and a language model according to the present disclosure demonstrates significantly improved performance compared to other language model integration methods. A key feature of this integration is the use of a fixed-size soft visual prompt regardless of the number of texts in the image, and in the test, the number of texts was set to 192. Unlike approaches where all OCR tokens are fed into the language model, the method of the present disclosure does not rely on unnecessarily large token lengths (indicated as |OCR|) for document information processing, thereby improving efficiency.

[0123] When considering actual applications, the approach can be evaluated even in scenarios where OCR and object detectors cannot be used. As can be seen in Table 4 below, the method / model according to the present disclosure can significantly improve performance compared to previous strategies. The neural network model of the present disclosure can effectively extract task-related information within a vision encoder in text-heavy tasks.

[0124] Model DocVQA ChartQA InfoVQA BLIP2-OPT-6.7B 3.7 4.6 11.0 BLIP2-FlanT5xxL-11B 8.6 4.4 11.4 LLaVA-Vicuna7B 5.5 0.5 2.4 LLaVA-Vicuna13B 5.9 1.4 3.1 Cream-Vicuna7B w / o Aux 45.8 50.0 22.8

[0125] Table 5 below shows the performance of the neural network model of the present disclosure when used independently. Recently released models are also included as subjects for comparison. The performance of the standalone neural network model of the present disclosure is slightly lower than that of specialized standalone models, but shows results similar to the latest visual document understanding models. In addition, it can be seen that performance improves when the language model is combined with the standalone neural network model. Tables 3 and 5 indicate that the language model can achieve a level of performance similar to state-of-the-art visual document understanding models even when relying solely on fixed-size visual prompts without directly observing images or OCR results.

[0126] Model Aux DocVQA ChartQA InfoVQA BROS o 68.1 - 24.8 Donut 67.5 41.8 21.7 Pix2Struct Base 72.1 56.0 38.2 Pix2Struct Large 76.6 58.6 40.0 LayoutLMv3 Base o 78.8 - - LayoutLMv3 Large o 83.4 - 45.1 UDOP o 84.7 - 47.4 Cream o 81.3 61.2 39.8

[0127] Overall, the experimental results indicate that the integration of the neural network model and the language model of the present disclosure is effective in the task of understanding visually text-rich images. By successfully integrating the strengths of visual understanding and language processing, the method of the present disclosure provides a powerful model that surpasses the integration of existing language models and demonstrates performance capable of competing with state-of-the-art standalone models across evaluated benchmarks.

[0128] The method described above may be provided as a computer program stored on a computer-readable recording medium for execution on a computer. The medium may continuously store a program executable by a computer, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or multiple hardware components combined, and may not be limited to a medium directly connected to a computer system but may exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Furthermore, other examples of media may include recording or storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.

[0129] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, these techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will understand that the various exemplary logical blocks, modules, circuits, and algorithmic steps described in connection with the disclosure herein may be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate such interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been generally described above in terms of their functional aspects. Whether such functions are implemented in hardware or in software depends on the design requirements imposed on the specific application and the overall system. Those skilled in the art may implement the functions described in various ways for each specific application, but such implementations should not be construed as departing from the scope of the present disclosure.

[0130] In a hardware implementation, the processing units used to perform the techniques may be implemented in one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described in this disclosure, computers, or a combination thereof.

[0131] Accordingly, the various exemplary logic blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed by any combination of general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors coupled with a DSP core, or any other combination of configurations.

[0132] In firmware and / or software implementations, techniques may be implemented as instructions stored on a computer-readable medium such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), magnetic or optical data storage devices, etc. The instructions may be executable by one or more processors, and may cause the processor(s) to perform specific aspects of the functions described in this disclosure.

[0133] When implemented in software, the techniques may be stored on a computer-readable medium as one or more instructions or code, or transmitted through a computer-readable medium. Computer-readable media include both computer storage media and communication media, including any medium that facilitates the transmission of a computer program from one place to another. Storage media may be any available media accessible by a computer. As a non-limiting example, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium accessible by a computer that can be used to transfer or store desired program code in the form of instructions or data structures. Additionally, any connection is appropriately made to the computer-readable medium.

[0134] For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair cable, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, coaxial cable, fiber optic cable, twisted pair cable, digital subscriber line, or wireless technologies such as infrared, radio, and microwave are included within the definition of a medium. As used herein, disk and disc include CD, laser disc, optical disc, DVD (digital versatile disc), floppy disk, and Blu-ray disc, wherein disks usually play data magnetically, whereas discs play data optically using a laser. The above combinations should also be included within the scope of computer-readable media.

[0135] The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other known form of storage medium. An exemplary storage medium may be connected to a processor so that the processor can read information from the storage medium or write information to the storage medium. Alternatively, the storage medium may be integrated into the processor. The processor and the storage medium may exist within an ASIC. The ASIC may exist within a user terminal. Alternatively, the processor and the storage medium may exist as separate components within the user terminal.

[0136] Although the embodiments described above have been described as utilizing aspects of the subject matter disclosed herein in one or more standalone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or a distributed computing environment. Furthermore, aspects of the subject matter in the present disclosure may be implemented in a plurality of processing chips or devices, and storage may be similarly affected across a plurality of devices. Such devices may include PCs, network servers, and portable devices.

[0137] Although the present disclosure has been described in relation to some embodiments, various modifications and changes may be made without departing from the scope of the present disclosure as understood by a person skilled in the art to which the invention of the present disclosure pertains. Furthermore, such modifications and changes should be considered to fall within the scope of the claims appended to this specification. Explanation of the symbols

[0138] 110: Image 120: Text prompt 130: Model 140: Answer

Claims

Claim 1 A method for enhancing image understanding, performed by at least one processor, comprising: receiving an image; receiving a text prompt; a first encoder dividing the image into a plurality of patches and generating a first set of embeddings based on the plurality of patches; and a second encoder generating a second set of embeddings based on a plurality of boxes containing at least one of text information or object information detected from the image. An image understanding enhancement method that is learned using a contrastive feature alignment technique, wherein the first encoder and the second encoder are trained by setting the specific box and the specific patch as associated data pairs (positive pairs) such that a specific box among the plurality of boxes detected as regions containing at least one of text information or object information within the image and a specific patch among the plurality of patches spatially corresponding to the position of the specific box are aligned at similar positions in a common feature space. Claim 2 delete Claim 3 A method for improving image understanding according to claim 1, wherein the aspect ratio of each of the plurality of patches is the same as the aspect ratio of the image. Claim 4 In claim 1, the step of generating the second set of embeddings comprises: the step of generating the first embedding based on first text information associated with an OCR (Optical Character Recognition) box detected from the image by the second encoder; and the step of generating the second embedding based on second text information associated with an object box detected from the image by the second encoder, wherein the second set of embeddings comprises the first embedding and the second embedding. Claim 5 In claim 4, the step of generating the second set of embeddings further comprises: a step in which a first detector detects an area containing text in the image as the OCR box; and a step in which a second detector detects an area containing a specific object in the image as the object box, wherein the first text information associated with the OCR box includes location information of the OCR box, recognized text information, and text information representing the first detector, and the second text information associated with the object box includes location information of the object box, recognized object label, and text information representing the second detector, a method for improving image understanding. Claim 6 A method for improving image understanding according to claim 1, wherein the step of generating the first output comprises concatenating the first set of embeddings generated by the first encoder and the second set of embeddings generated by the second encoder in a sequence direction and inputting them to the cross attention layer of the decoder. Claim 7 delete Claim 8 delete Claim 9 In claim 4, the step of generating the first set of embeddings comprises: dividing the image into the plurality of patches—the plurality of patches including the first patch containing the OCR box and the second patch containing the object box—the first encoder generating a third embedding associated with the first patch; and the first encoder generating a fourth embedding associated with the second patch, wherein the first embedding and the third embedding are aligned at a first position in a common feature space, and the second embedding and the fourth embedding are aligned at a second position in the common feature space, a method for improving image understanding. Claim 10 A method for improving image understanding according to claim 1, wherein the text prompt is a natural language question associated with the image, and the method further comprises the step of generating a token sequence associated with an answer to the natural language question referencing the image based on the first output. Claim 11 In claim 10, the above decoder uses an attention mechanism having a unidirectional attention flow, a method for improving image understanding. Claim 12 The method for enhancing image understanding according to claim 1, wherein the step of generating the first output comprises the step of the decoder generating the first output based on the first set of embeddings, the second set of embeddings, the text prompt, and the learned query, and further comprises the step of generating a soft visual prompt based on the first output; and the step of a language model generating a second output based on the soft visual prompt and the text prompt. Claim 13 In claim 12, an image understanding enhancement method in which the number of vectors included in the learned query and the number of vectors included in the soft visual prompt are the same. Claim 14 In claim 12, the image understanding enhancement method, wherein the decoder uses an attention mechanism having a bi-directional attention flow. Claim 15 In claim 12, an image understanding enhancement method in which the decoder is updated with additional learning while the first encoder, the second encoder, and the language model are in a frozen state. Claim 16 A computer program stored on a computer-readable recording medium for executing a method according to any one of paragraphs 1, 3 through 6, and 9 through 15 on a computer. Claim 17 As a system, communication module; memory; and includes at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, wherein the at least one program includes instructions for receiving an image and receiving a text prompt, wherein a first encoder divides the image into a plurality of patches and generates a first set of embeddings based on the plurality of patches, wherein a second encoder generates a second set of embeddings in a plurality of boxes containing at least one of text information or object information detected from the image, and wherein a decoder generates a first output based on the first set of embeddings, the second set of embeddings, and the text prompt, wherein the first encoder and the second encoder perform contrast feature alignment such that a specific box among the plurality of boxes detected as an area containing at least one of text information or object information within the image and a specific patch among the plurality of patches spatially corresponding to the location of the specific box are aligned at similar positions in a common feature space, thereby setting the specific box and the specific patch as associated data pairs (positive pairs). A system trained using a contrastive feature alignment technique.