Electronic device and method thereof for restoring images using intrinsic information of intermediate layer within model trained to output extrinsic information
The image restoration model enhances image resolution and text recognition accuracy by using an encoder, sub-model, and decoder structure to combine intrinsic information, effectively addressing the challenge of restoring distorted or low-resolution images.
Patent Information
- Application Number
- JP2025032619
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-18
- Filing Date
- 2025-03-03
- Publication Date
- 2025-09-18
AI Technical Summary
Existing image processing technologies struggle to accurately restore images with distorted or low-resolution text, such as license plates, due to limitations in recognizing and enhancing characters effectively.
An electronic device employs an image restoration model that includes an encoder for feature extraction, a sub-model for determining a text probability map, a synthesis layer for combining intrinsic information from intermediate layers, and a decoder to generate high-resolution images, utilizing a neural network structure like TRBA-ResNet-BiLSTM-attention mechanism to enhance image resolution and accuracy.
The model effectively enhances image resolution and accuracy of text recognition, reducing distortions and improving the visibility of characters, particularly in low-resolution images, by leveraging intrinsic information from intermediate layers.
Smart Images

Figure 2025135580000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to an electronic device and method for image reconstruction using implicit information from intermediate layers in a model trained to output explicit information. [Background technology]
[0002] Techniques have been developed to process photographs and / or videos using artificial intelligence, for example, to classify objects (e.g., people, animals, and / or vehicles) captured in the photographs and / or videos, or to recognize one or more characters (or strings of characters) associated with the photographs and / or videos.
[0003] The preceding information may be provided as related art to aid in the understanding of the present disclosure. No assertion or determination is being made as to the applicability of any of the preceding as prior art pertaining to the present disclosure. Summary of the Invention [Means for solving the problem]
[0004] According to one embodiment, a non-transitory computer-readable storage medium may be provided that includes instructions that, when executed individually or collectively by at least one processor of an electronic device, can cause the electronic device to receive a request to restore a first image at a first resolution to an image at a second resolution greater than the first resolution. When executed individually or collectively by the at least one processor, the instructions can cause the electronic device to execute, based on the received request, an image restoration model that includes an encoder for extracting feature information from the first image, a sub-model for determining a text probability map for the first image, a synthesis layer disposed before an output layer trained to output the text probability map for combining intrinsic information of an intermediate layer of the sub-model with the feature information, and a decoder coupled to the synthesis layer for generating the image at the second resolution. When executed individually or collectively by the at least one processor, the instructions can cause the electronic device to provide, in response to the request, a second image at the second resolution obtained based on the execution of the image restoration model.
[0005] According to one embodiment, an electronic device may include a memory storing instructions and at least one processor configured to execute the instructions. When executed individually or collectively by the at least one processor, the instructions may cause the electronic device to receive a request to restore a first image at a first resolution to an image at a second resolution greater than the first resolution. When executed individually or collectively by the at least one processor, the instructions may cause the electronic device to execute, based on the received request, an image restoration model including: an encoder for extracting feature information from the first image; a sub-model for determining a text probability map for the first image; a synthesis layer disposed before an output layer trained to output the text probability map for combining the feature information with intrinsic information of an intermediate layer of the sub-model; and a decoder coupled to the synthesis layer for generating the image at the second resolution. When executed individually or collectively by the at least one processor, the instructions may cause the electronic device to provide, in response to the request, a second image at the second resolution obtained based on the execution of the image restoration model.
[0006] According to one embodiment, a method for an electronic device may be provided. The method may include, based on receiving an image, obtaining a sub-model trained to output a text probability map representing one or more characters associated with the image. The method may include using the sub-model to train an image restoration model including an encoder for extracting feature information from an input image, a synthesis layer receiving the input image and combining the feature information with intrinsic information from an intermediate layer preceding an output layer of the sub-model, and a decoder coupled to the synthesis layer for generating an output image having a second resolution greater than a first resolution of the input image. The method may include providing the image restoration model as part of a software application for restoring images.
[0007] According to one embodiment, an electronic device may include a memory storing instructions and at least one processor configured to execute the instructions. When executed individually or collectively by the at least one processor, the instructions may cause the electronic device to obtain a sub-model trained to output a text probability map representing one or more characters associated with an image based on receiving the image. When executed individually or collectively by the at least one processor, the instructions may cause the electronic device to use the sub-model to train an image restoration model including an encoder for extracting feature information from an input image, a synthesis layer receiving the input image and combining the feature information with intrinsic information from an intermediate layer preceding an output layer of the sub-model, and a decoder coupled to the synthesis layer for generating an output image having a second resolution greater than a first resolution of the input image. When executed individually or collectively by the at least one processor, the instructions may cause the electronic device to provide the image restoration model as part of a software application for restoring images. [Brief explanation of the drawings]
[0008] [Figure 1] 1 illustrates an exemplary block diagram of an electronic device for restoring at least a portion of an image. [Figure 2] 1 illustrates an exemplary block diagram of the structure of an image restoration model executed by an electronic device, according to one embodiment. [Figure 3] 1 illustrates an exemplary block diagram of the structure of a sub-model within an image restoration model trained to output a text probability map. [Figure 4] 1 illustrates an exemplary block diagram of the structure of an image restoration model executed by an electronic device, according to one embodiment. [Figure 5]1 illustrates an exemplary block diagram of an image restoration model associated with a teacher model. [Figure 6] 1 shows a graph illustrating the performance of an electronic device implementing an image restoration model, according to one embodiment. [Figure 7] 1 shows a graph illustrating the performance of an electronic device implementing an image restoration model, according to one embodiment. [Figure 8] [Figure 8a] shows at least one license plate (license plate or number plate) that is a subject included in an image restored by an image restoration model according to one embodiment. [Figure 8b] shows at least one license plate (license plate or number plate) that is a subject included in an image restored by an image restoration model according to one embodiment. [Figure 9] FIG. 1 is a diagram for explaining an overconfidence phenomenon. DETAILED DESCRIPTION OF THE INVENTION
[0009] Various embodiments of the present document will now be described with reference to the accompanying drawings.
[0010] 1 shows an example block diagram of an electronic device 101 for restoring at least a portion of an image 150. The electronic device 101 may be configured to at least partially restore or enhance the image 150. Restoring or enhancing the image 150 may include operations that compensate for distortions in the image 150, such as blur, persistence of vision, optical flow, etc., to improve the visibility of objects represented by the image 150.
[0011] Referring to FIG. 1 , an image 150 including a portion 152 associated with a license plate (or number plate) is illustratively shown. For example, the image 150 can be transmitted from an external electronic device to the electronic device 101 via the communication circuitry 130. For example, the image 150 can be acquired using a camera 140 included in the electronic device 101. For example, the image 150 can be a file formatted based on a format for compressing and storing digital images, such as JPEG (Joint Photographic Experts Group) or PNG (Portable Network Graphics). For example, the image 150 can include raw data acquired from the camera 140. For example, the image 150 can be included in a sequence of image frames (e.g., a video) that are included in a video and configured to be displayed sequentially. The means for acquiring or receiving the image 150 is not limited to the communication circuitry 130 and / or the camera 140 shown in FIG. 1 .
[0012] 1 , an exemplary object, such as a vehicle, may be captured. Depending on the environment in which the object is captured, image 150 may be distorted. For example, if the object moves (e.g., a vehicle moving) and / or a camera (e.g., camera 140) controlled to capture image 150 moves (or shakes), the appearance of the object represented by the pixels of image 150 may be distorted. According to one embodiment, electronic device 101 may at least partially reduce or remove distortions occurring in image 150 to sharpen the appearance of the object represented by image 150.
[0013] Referring to FIG. 1 , an exemplary hardware configuration of an electronic device 101 for at least partially restoring an image 150 is shown. For example, the electronic device 101 may include a personal computer (PC), such as a laptop or desktop, a smartphone, a smartpad, or a tablet PC. For example, the electronic device 101 may include a smart accessory, such as a smartwatch, a smart ring, and / or a head-mounted device (HMD). For example, the electronic device 101 may be referred to as a mobile device, a user equipment (UE), a multifunction device, a portable communication device, and / or a handheld device. For example, the electronic device 101 may be included as an electronic control unit (ECU) in a vehicle (e.g., an electric vehicle (EV)). For example, the electronic device 101 may include a server of a service provider that provides a service for restoring the image 150. The server may include one or more PCs and / or workstations.
[0014] 1 , an electronic device 101 according to one embodiment may include at least one of a processor 110, a memory 120, a communication circuit 130, or a camera 140. According to an embodiment, the communication circuit 130 and / or the camera 140 may not be included in the electronic device 101. For example, the communication circuit 130 and / or the camera 140 may be located external to the electronic device 101 and electrically connected to the electronic device 101.
[0015] Referring to FIG. 1 , the processor 110, the memory 120, the communication circuit 130, and the camera 140 may be electrically and / or operably coupled with each other by electronic components such as a communication bus 102. Hereinafter, the term "operably coupled" may refer to a direct or indirect connection between a first electronic component and a second electronic component, established by wire or wirelessly, such that the first electronic component controls the second electronic component. Although illustrated based on different blocks, embodiments are not limited thereto, and some of the electronic components in FIG. 1 (e.g., at least a portion of the processor 110, the memory 120, and the communication circuit 130) may be included in a single integrated circuit, such as a system on a chip (SoC). The types and / or number of electronic components included in the electronic device 101 are not limited to those illustrated in FIG. 1 . For example, the electronic device 101 may include only some of the electronic components illustrated in FIG. 1 .
[0016] According to one embodiment, the processor 110 of the electronic device 101 may include circuitry (e.g., processing circuitry) for processing data based on one or more instructions. The circuitry for processing data may include, for example, an arithmetic and logic unit (ALU), a floating point unit (FPU), a field programmable gate array (FPGA), a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), and / or an application processor (AP). For example, the number of processors 110 may be one or more. The processing circuitry of the processor 110 that loads (or fetches) instructions and performs calculations corresponding to the loaded instructions may be referred to as or refer to a core circuit (or core). For example, the processor 110 may have a multi-core processor architecture including multiple core circuits, such as a dual-core, quad-core, hexa-core, or octa-core. The functions and / or operations described with reference to the present disclosure may be performed individually and / or collectively by one or more processing circuits included in processor 110.
[0017] According to one embodiment, memory 120 of electronic device 101 may include circuitry for storing data and / or instructions input and / or output to processor 110. Memory 120 may include, for example, volatile memory, such as random-access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM). Non-volatile memory may be referred to as storage. Volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). Non-volatile memory may include, for example, at least one of programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, a hard disk, a compact disk, a solid state drive (SSD), and an embedded multimedia card (eMMC). Memory 120 may include one or more storage media (e.g., the volatile and / or non-volatile memory described above) arranged in a distributed manner throughout electronic device 101. Processor 110 of electronic device 101 may execute instructions from memory 120 within electronic device 101 to perform the functions and / or operations represented by the instructions. For example, if electronic device 101 includes at least one processor, the at least one processor may be configured to execute the instructions collectively or individually.
[0018] According to one embodiment, the communications circuitry 130 of the electronic device 101 may include hardware for supporting transmission and / or reception of electrical signals between the electronic device 101 and an external electronic device (e.g., a user terminal configured to transmit the image 150). The communications circuitry 130 may include, for example, at least one of a modem, an antenna, and an optical / electronic (O / E) converter. The communications circuitry 130 may support transmission and / or reception of electrical signals based on various types of protocols, such as Ethernet, a local area network (LAN), a wide area network (WAN), wireless fidelity (WiFi), near field communication (NFC), Bluetooth, Bluetooth low energy (BLE), Zigbee, long term evolution (LTE), fifth generation (5G), new radio (NR), sixth generation (6G), and / or above-6G.
[0019] According to one embodiment, the camera 140 of the electronic device 101 may include one or more light sensors (e.g., a charged coupled device (CCD) sensor, a complementary metal oxide semiconductor (CMOS) sensor) that generate electrical signals indicative of the color and / or brightness of light. The multiple light sensors included in the camera 140 may be arranged in the form of a two-dimensional grid. The camera 140 may acquire the electrical signals of each of the multiple light sensors substantially simultaneously to generate two-dimensional frame data corresponding to the light reaching the light sensors of the two-dimensional grid. For example, photographic data captured using the camera 140 may refer to one two-dimensional frame data acquired from the camera 140. For example, video data captured using the camera 140 may refer to a sequence of two-dimensional frame data acquired from the camera 140.
[0020] 1 , according to one embodiment, the processor 110 of the electronic device 101 can execute an image restoration program 125 to at least partially restore or enhance an image 150. The processor 110 (e.g., a CPU, a GPU, and / or an NPU) that executed the image restoration program 125 can perform calculations to restore the image 150. The calculations may involve a computational model (e.g., an artificial neural network and / or a neural network) configured to mimic neural activity of a living organism. The neural activity may include, for example, cognitive activity, reasoning activity, and / or creative activity of the living organism. For example, instructions representing the computational model, mathematical formulas associated with the computational model, and / or constants (e.g., coefficients and / or weights) included in the mathematical formulas may be at least partially included in the image restoration program 125.
[0021] According to one embodiment, the processor 110 of the electronic device 101 can restore or enhance a portion 152 of the image 150 in which at least one character is captured (e.g., a portion in which an object with one or more characters printed thereon, such as a license plate and / or a cover plate, is captured). For example, within the image 150, the electronic device 101 can extract or segment or crop the portion 152 associated with the at least one character. The portion 152 may be referred to as a region of interest (ROI). The processor 110 can execute the image restoration program 125 to restore or enhance the portion 152.
[0022] In one embodiment, electronic device 101 can recognize text associated with a scene, such as image 150 (e.g., text that indicates that the scene is captured or contained within the scene), and increase or enhance the resolution of the scene. For example, if electronic device 101 detects one or more characters from a scene with a relatively low resolution (or small size), electronic device 101 can use the form and / or appearance of the detected one or more characters to generate another scene that corresponds to the scene and has a higher resolution (or larger size) than the resolution of the scene. For example, from a scene with width w and height h, relative to a scaling factor f, electronic device 101 can generate or output a scene with width fw and height fh.
[0023] In one embodiment, the image restoration program 125 and / or the artificial intelligence driven by the image restoration program 125, from the perspective of recognizing text and generating high-resolution scenes, may be referred to as scene text image super-resolution (STISR) and / or a model for STISR. The performance of STISR may be evaluated using the accuracy of characters included in high-resolution scenes generated by executing STISR (e.g., STISR accuracy).
[0024] 1 , image 160 is shown as an output result of electronic device 101 restoring portion 152 of image 150. Image 150 and / or portion 152 may be referred to as an input image in terms of being input to processor 110 of electronic device 101. Image 160 may be referred to as an output image in terms of output data corresponding to the input image. According to one embodiment, electronic device 101 may obtain information representing one or more characters associated with portion 152 using an artificial intelligence model trained to recognize one or more characters from an image. Using the information, electronic device 101 may generate or output image 160 as a high-resolution image corresponding to portion 152.
[0025] 1 , image 160 can have a larger size and / or a higher resolution than portion 152. The dimensions (e.g., width and / or height) of image 160 can be larger than the dimensions of portion 152. For example, image 160 can have dimensions and / or resolution such as image 150. In one embodiment, electronic device 101 receives image 150 and / or portion 152 from an external electronic device via communications circuitry 130. Electronic device 101 can receive a request to restore portion 152 of image 150, which has a first resolution, to image 160, which has a second resolution greater than the first resolution. Electronic device 101 can identify or detect image 150 and / or portion 152 from a signal received from the external electronic device. The signal can include a command and / or operand indicating a request to restore portion 152. In one embodiment, where the entire image 150 including the portion 152 is received, the processor 110 of the electronic device 101 can extract or segment the portion 152 in which one or more characters associated with the subject, such as a license plate, are captured. The portion 152 can be used as the image used for reconstruction.
[0026] Based on a request to restore image 150 and / or portion 152, electronic device 101 may execute an artificial intelligence model (e.g., an image restoration model) provided by image restoration program 125. In response to the request, electronic device 101 may provide image 160 at the second resolution obtained based on execution of the image restoration model. For example, electronic device 101 may transmit a signal including image 160 to an external electronic device via communication circuitry 130.
[0027] In one embodiment, the image restoration model executed by the image restoration program 125 may include a sub-model trained to recognize one or more characters associated with (e.g., appearing to be captured by) an input image (e.g., portion 152 and / or image 150 including portion 152) input to the image restoration model. The sub-model may be trained to output information (e.g., explicit information) readable by the processor 110 executing a software application distinct from the image restoration model and / or image restoration program 125, representing one or more characters associated with the input image, the degree to which each of the one or more characters is associated with the input image (e.g., the probability that the one or more characters were captured by the input image), and / or the positional relationship of the one or more characters (e.g., the position and / or order of each of the one or more characters within a string of characters).
[0028] For example, the information output from the sub-models may be referred to as text probability information, in that it includes probabilities indicating text that appears captured by the input image. The text probability information may also be referred to as text categorical information, text probabilities, text probability maps, text prior information, and / or text distributions. For example, the text probability information may include text categorical information and / or information indicating visual cues for text in an image.
[0029] According to one embodiment, the electronic device 101 can be trained to generate an image 160 using intermediate states and / or intermediate information of a sub-model trained to output extrinsic information such as text probability information. For example, among the nodes (e.g., perceptrons) of the sub-model separated by multiple layers, values of nodes in the output layer, including nodes corresponding to each element of the text probability information, and other nodes, may be directly transmitted to other sub-models of the image restoration model. For example, intermediate layers of a sub-model can be connected to other sub-models of the image restoration model.
[0030] For example, the values of nodes included in the intermediate layer may be implicit information, which is distinguished from extrinsic information. The implicit information may include more information about the input image than text probability information, which includes only the probability that the input image (e.g., portion 152 and / or image 150) corresponds to each of multiple characters. By executing an image restoration model using the implicit information, electronic device 101 may more accurately restore portion 152. For example, electronic device 101 may obtain or generate image 160 that more accurately represents one or more characters included in portion 152. In the example above, upon receiving a request to repeatedly restore portion 152, multiple images (e.g., image 160) generated in response to the request may contain similar characters because they more accurately recognize or represent one or more characters from portion 152.
[0031] Hereinafter, an exemplary structure of the image restoration model executed by the image restoration program 125 and a process of training the image restoration model will be exemplarily described with reference to FIGS.
[0032] 2 illustrates an exemplary block diagram of the structure of an image restoration model executed by an electronic device (e.g., electronic device 101 of FIG. 1), according to one embodiment. Electronic device 101 and / or processor 110 of FIG. 1 can execute image restoration program 125 to implement the image restoration model described with reference to FIG. 2.
[0033] Hereinafter, an operation of executing an artificial intelligence model, such as an image restoration model, may include an operation of performing one or more calculations associated with the artificial intelligence model using a processor of an electronic device (e.g., processor 110 of FIG. 1 , which includes a GPU and / or an NPU). The operation of executing an artificial intelligence model may include an operation of inputting commands (or instructions) representing the calculations to a GPU and / or an NPU for execution of the calculations by the GPU and / or the NPU. The operation of executing an artificial intelligence model may include an operation of inputting data (e.g., an input image, such as image 150 and / or portion 152 of FIG. 1 ) to be at least partially modified by the calculations to a GPU and / or an NPU. While an operation of executing an artificial intelligence model based on a GPU and / or an NPU has been described as an example, embodiments are not limited thereto, and an operation of executing an artificial intelligence model using a CPU may also be performed similarly to the operations described above.
[0034] Referring to Figure 2, calculations performed by the image restoration model are shown as multiple blocks to distinguish the type and / or order of the calculations. Any block in Figure 2 may correspond to a group of calculations performed during execution of an artificial intelligence model (e.g., an image restoration model). Each block in Figure 2 may be referred to as an operation, layer, sub-model, and / or module for the artificial intelligence model. Referring to Figure 2, an exemplary image restoration model is shown, including a teacher model 210 connected to the image restoration model to train at least a portion of the image restoration model.
[0035] For example, the image restoration model may include an encoder (e.g., a combination of spatial transformer networks (STN) operations 241 and convolution operations 242) for extracting feature information from an image. The encoder including the STN operations 241 and / or convolution operations 242 may include a shallow convolutional neural network (CNN), which reduces the loss of structural information (or spatial information) required for image restoration. The encoder of the image restoration model (or STISR) may include a relatively small number of layers to reduce the loss of structural information (or spatial information) of the low-resolution image when extracting features of the low-resolution image to perform a low-level vision task (e.g., a task of increasing the resolution of the image). By executing the encoder of the image restoration model, the electronic device may generate or obtain feature information of the input image 202. The feature information may include summarized (or reduced-dimensional) information of the input image 202 for identifying or distinguishing the input image 202. The feature information may include the locations and / or characteristics of one or more pixels uniquely contained in the input image 202, such as feature points (or key points) and / or boundary lines.
[0036] For example, the image restoration model may include a sub-model 220 for determining a text probability map for the input image 202. The teacher model 210 may use knowledge distillation to generate training information (e.g., ground truth data and input data corresponding to the ground truth data) used to learn the sub-model 220. The calculations of the sub-model 220 and the number of parameters (e.g., coefficients and / or weights) used in these calculations may be fewer than the calculations of the teacher model 210 and the number of parameters used in the calculations of the teacher model 210. For example, the sub-model 220 may be pretrained by the teacher model 210 running with more parameters than the parameters for the sub-model 220.
[0037] In one embodiment, the teacher model 210 used to train the sub-models 220 may be trained to recognize one or more characters from a scene, such as the image 201. From the perspective of recognizing characters, the teacher model 210 may be referred to as a scene-text recognizer (STR) and / or STR model. The teacher model 210 may be configured to recognize or process features, such as the shape and / or location of one or more characters in the image 201.
[0038] 2, the types and orders of calculations of the teacher model 210 and the sub-model 220 may be similar or consistent with each other. For example, when executing the sub-model 220, the electronic device may sequentially execute an encoding operation 220a, a sequence modeling operation 220b, a decoding prediction operation 220c, and a linearization operation 220d on the input image 202 to obtain or generate output data (e.g., text probability information and / or a text probability map). The operations (e.g., the encoding operation 220a, the sequence modeling operation 220b, the decoding prediction operation 220c, and the linearization operation 220d) sequentially executed in the sub-model 220 may correspond to the operations (e.g., the encoding operation 210a, the sequence modeling operation 210b, the decoding prediction operation 210c, and the linearization operation 210d) sequentially executed in the teacher model 210, respectively. The connections of the above operations may have a TRBA (thin plate spline transformation (TPS)-Residual neural network (ResNet)-BiLSTM (bidirectional long-short term memory)-attention mechanism) structure. An exemplary structure of the sub-model 220 having the TRBA structure is described in detail with reference to FIG. 3 . The embodiment is not limited thereto, and the structure of the sub-model 220 may adopt other structures (or topologies), such as a convolution-recurrent neural network (CRNN), an autonomous, bidirectional and iterative network (ABINet), and / or a permuted autoregressive sequence (PARseq). The output layer of the sub-model 220 may include values determined by a calculation performed for the linearization operation 220d. The values included in the output layer may be text probability information.
[0039] According to one embodiment, an electronic device can train sub-model 220 using teacher model 210, to which image 201 having a relatively high resolution is input. For example, an electronic device executing teacher model 210 can determine, from image 201, a text probability map representing one or more characters associated with image 201. The electronic device can train sub-model 220 using other images having a lower resolution than image 201 and the determined text probability map.
[0040] Referring to FIG. 2 , the output layer of the sub-model 220 may be associated with a linearization operation 220d. Within the sub-model 220, intrinsic information used in the linearization operation 220d, including the result of executing the decoding prediction operation 220c (or the state of any intermediate layer for the decoding prediction operation 220c), may be provided to a synthesis layer 243. Before being provided to the synthesis layer 243, the intrinsic information may be input to a projection model 230. Using the projection model 230, the electronic device may sequentially perform a linearization operation 232a, a PReLU (Parametric rectified linear unit) operation 232b (e.g., an operation included in the sub-model 232), and a prior interpreter operation 234 on the intrinsic information. The intrinsic information, at least partially modified by the projection model 230, may be input to the synthesis layer 243. The projection model 230 may be referred to as a non-categorical prior (NCAP) because it outputs information inherent to the sub-models 220 trained to generate categorical information (e.g., non-categorical information). The combination of the sub-models 220 and the projection model 230 may be referred to as a scene-text recognizer (STR). The information output by the projection model 230 (e.g., information sent from the projection model 230 to the synthesis layer 243) may be referred to as prior knowledge information.
[0041] The combination of the sub-model 220 and the projection model 230 allows an electronic device executing the image restoration model to generate an output image 203 using text information (e.g., text probability information) inferred from the input image 202. The encoder, which is a combination of an STN (spatial transformer networks) operation 241 and a convolution operation 242, allows an electronic device executing the image restoration model to generate an output image 203 using non-text information (e.g., structural information) inferred from the input image 202. In terms of using both text information and non-text information, the image restoration model may be a model that supports multimodality.
[0042] 2, a synthesis layer 243 may be configured to combine the feature information with implicit information from the intermediate layers of the sub-models, which are placed before an output layer trained to output a text probability map. For example, the electronic device can perform the calculations represented by the synthesis layer 243 using all of the feature information, including the results of executing the convolution operation 242 of the encoder, and the text probability maps output or generated from the sub-models 220 and / or the projection model 230.
[0043] 2, the image restoration model may use information generated by a synthesis layer 243 to perform decoder operations 244 to generate an output image 203 having a resolution greater than that of the input image 202. The decoder operations 244 may be trained to use the information generated by the synthesis layer 243 to generate an output image 203 that has a resolution greater than that of the input image 202 and / or a size greater than that of the input image 202 and that is related to (e.g., includes the content of) the input image 202. The output image 203 may be provided as a result of restoring or enhancing the input image 202.
[0044] As described above, according to one embodiment, an electronic device can generate or obtain an output image 203 from an input image 202 by executing an image restoration model including a sub-model 220 trained to output one or more characters appearing to be captured by the input image 202 and a text probability map indicating the location of the one or more characters. The image restoration model can include a synthesis layer 243 connected (e.g., indirectly connected via a projection model 230) to an intermediate layer of the sub-model 220 (e.g., an intermediate layer for performing the decoding prediction operation 220c) to extract the intrinsic information used to determine the text probability map, which is extrinsic information. For example, to reduce or prevent distortion of the output image 203 due to errors that may be contained in the text probability map (e.g., a result of incorrectly recognizing at least one character from the input image 202), the electronic device can synthesize or generate the output image 203 from the text probability map using the intrinsic information used to determine the text probability map, including various information about the input image 202.
[0045] In one embodiment, because the intrinsic information includes higher-dimensional information than the text probability map, the electronic device can effectively resolve domain gaps due to differences in resolution between the input image 202 and the output image 203. For example, without domain transfer, the electronic device can obtain or generate information (e.g., the intrinsic information) used to reduce or eliminate the domain gap.
[0046] In one embodiment, to reconstruct the output image 203 from the input image 202, a sub-model 220 included in the image restoration model can be retrained by the teacher model 210 to obtain information about one or more characters (e.g., a text probability map). The retrained sub-model 220 can generate or output features (e.g., discriminative features) useful to the remaining layers of the image restoration model (e.g., the projection model 230, the synthesis layer 243, and / or the decoder 244) that are connected to the sub-model 220 and execute by generating the output image 203. The image restoration model including the retrained sub-model 220 can be trained using ground truth data (e.g., a pair of the output image 203 and an input image 202 obtained by distorting the output image 203 and having a smaller resolution and / or a smaller size than the output image 203).
[0047] For example, an image restoration model can be trained to output an output image 203 as a result of enhancing an input image 202 through a training process that includes a first step of retraining a pre-trained sub-model 220 and a second step of training an image restoration model including the retrained sub-model 220. The first step of the training process is described with reference to FIG. 3. The second step of the training process is described with reference to FIG. 4.
[0048] 3 shows an example block diagram of the structure of a sub-model 220 in an image restoration model trained to output a text probability map. The electronic device 101 and / or the processor 110 of FIG. 1 can execute the image restoration program 125 to obtain or execute the sub-model 220 and / or the image restoration model including the sub-model 220 described with reference to FIG. 3.
[0049] According to one embodiment, an electronic device may obtain a sub-model 220 trained to output a text probability map representing one or more characters associated with an image based on receiving the image. The electronic device may re-train (e.g., fine-tune) the obtained sub-model 220 using a loss function. The loss function may be configured or defined to generate not only extrinsic information (e.g., text probability information) output from the sub-model 220, but also intrinsic information representing discriminative features used by an image restoration model including the sub-model 220.
[0050] 3, an exemplary structure of the submodel 220 having a TRBA structure is shown. Based on the TRBA structure, the submodel 220 may include a TPS model 310 for TPS operation, a backbone model 320 (e.g., ResNET), a BiLSTM 330, a first feedforward model 340, multiple RNN decoders 350, and a second feedforward model 360, all of which are concatenated. The electronic device 101 may input an image 301 (x∈R̂(h×w×c))) at the input layer of the TPS model 310. The x in x∈R̂(h×w×c) may represent an input image (e.g., image 301) input to the input layer. For example, the image 301 may have a height h and a width w and c channels (e.g., three channels each for red, green, and blue).
[0051] In one embodiment, the sub-model 220 may be trained (eg, pre-trained) to output the extrinsic information p of equation (1).
[0052]
number
[0053] The extrinsic information p in Equation 1 is the output data (e.g., P0, P1, ..., P) of the second feedforward model 360 in FIG. t ) The output data P of the second feedforward model 360 can be t (where 0≦t≦l) is the output data h=[h1, h2, ..., h3] of the RNN decoder 350 from the first timing (or the first time step) to the tth timing (or the tth time step). i The extrinsic information p of Equation 1 can be determined based on the projection of the output data h=[h1, h2, ..., hi] of the RNN decoder 350. For example, the output data of the sub-model 220 can be set as shown in Equation 2.
[0054]
number
[0055] h in Math 2 i may be the intermediate state vector of the hidden layer (e.g., the RNN decoder 350) of the submodel 220 at the tth timing (or the Tth time step).
[0056] In one embodiment, the sub-model 220 may be trained using a loss function to increase and / or maximize the probability difference and / or margin between classes determined by the sub-model 220, as well as the cross-entropy loss. The loss function may be defined to generate discriminative features for the classes of the sub-model 220 (e.g., classes corresponding to each of multiple characters) and / or to alleviate confusion between the classes. The loss function may be used to retrain a pre-trained sub-model 220 to output extrinsic information p from the image 301.
[0057] For example, the loss function (L_str) set for retraining the submodel 220 may be set to increase or maximize the distance and / or margin from the crystal boundary. For example, it can be defined to maximize the margin of the decision boundary between a specific class y_i and another class y_(j,j≠i). For example, the loss function ( ) can be defined as the sum of L_rec in Equation 3 and L_aux in Equation 4 (L_str=L_rec+L_aux).
[0058]
number
[0059]
number
[0060] L_rec in Equation 3 is the cross-entropy for training the submodel 220, and p t、i may represent the probability of matching the i-th class among the classes (e.g., characters) classified by the sub-model 220. |A| in the number 3 may represent the total number of classes. l in the number 3 may correspond to the number of RNN decoders 350.
[0061] L_aux in Equation 4 may be defined to obtain discriminative features (or intrinsic information) using the sub-model 220. p_(y_t) in Equation 4 may represent the output probability of the sub-model 220 corresponding to a character that appears as a correct answer according to truth data when the sub-model 220 is trained using truth data representing the image 301 and one or more characters contained in the image 301. Referring to Equation 3, pt,i may represent the output probability of the sub-model 220 corresponding to the incorrect character, since i does not match a character that appears as a correct answer (i ≠ p_(y_t)). ε in Equation 4 is a real number (e.g., 10) defined so that the result value of the log function does not decrease to a value that is too small (e.g., negative infinity). -7) The l in numbers 3 and 4 can represent the length of the maximum string, and the i in numbers 3 and 4 can be a variable that changes within the number of the entire class (|A|).
[0062] Using a loss function of the difference between the output data of the submodel 220 and the truth data for the image 301 (e.g., L_str=L_rec+L_aux), the submodel 220 calculates the output data h=[h1, h2, ... h i ] may be trained to have discriminative features. The output data h=[h1, h2, ..., h i ] can be used to execute and / or train an image restoration model including the sub-model 220. Hereinafter, with reference to FIG. 3, the output data h=[h1, h2, ..., h i An exemplary structure of an image restoration model using [mathematical formula - see original document] is described below.
[0063] 4 shows an exemplary block diagram of the structure of an image restoration model executed by an electronic device, according to one embodiment. The electronic device 101 and / or processor 110 of FIG. 1 can execute the image restoration program 125 to implement or train the image restoration model described with reference to FIG. 4.
[0064] Referring to FIG. 4, the image restoration model may include a TPS model 420 and a shallow CNN 421. By performing calculations represented by the TPS model 420 and the shallow CNN 421 from the input image 402, the electronic device 101 may extract low-level feature information. By combining the feature information with position embedding data for a synthesis operation, the electronic device 101 may obtain feature information of [F_v∈R]^(c×hw). C in R^(c×hw) represents the dimension of the feature information and may correspond to the dimension of the information output from the output layer of the shallow CNN 421. hw in R^(c×hw) may represent the size (e.g., the number of parameters arranged in one dimension) of the flattened information (e.g., one-dimensional information) of the input image 402.
[0065] By performing the calculations represented by the TPS model 420, the electronic device 101 can adjust the shape of characters in the input image 402 so that the characters have a uniform shape. For example, the information output from the Flatten model 422 connected to the shallow CNN 421 may correspond to F_v in equation (5).
[0066]
number
[0067] x_LR in Equation 5 may represent an input image 402 having a relatively low resolution. PE in Equation 5 may represent position embedding data combined with feature information. Flatten in Equation 5 may represent an operation for converting multidimensional information into one-dimensional information. Enc1 in Equation 5 may represent an operation performed by the shallow CNN 421. According to one embodiment, the image restoration model may be trained to use information indicating spatial characteristics of the image (e.g., position embedding data PE in Equation 5) to consider distances between pixels in the image while calculating feature information.
[0068] When processing the input image 402 using the image restoration model, the electronic device 101 may perform a first operation of processing the input image 402 using the TPS 420 and / or the shallow CNN 421 in parallel (or substantially simultaneously) and a second operation of processing the input image 402 using the sub-model 220. The first operation and the second operation may be performed substantially simultaneously by different processors included in the electronic device 101. Using the sub-model 220 in a state trained based on the operation described with reference to FIG. 3, the electronic device 101 may obtain intrinsic information p_NCAP∈R̂(l×embed) from the input image 402. The intrinsic information p_NCAP may be determined or calculated based on Equation 6.
[0069]
number
[0070] The term STR in Equation 6 means scene text recognizer, and STR stu、dec STR may represent the operations performed in the decoder of the submodel 220 (for example, the BiLSTM, Attention mechanism, and Linear groups in the submodel 220). star、encmay represent the operations performed in the encoder of the submodel 220 (e.g., the ResNet in the submodel 220). LR may represent the input image 402 having a relatively low resolution.
[0071] Using the information obtained from the NCAP projector 410 (e.g., p_NCAP in Equation 6), the electronic device derives the feature information F in Equation 7 from the projection model 230. p can be obtained or calculated.
[0072]
number
[0073] By performing a softmax operation and / or a layer normalization operation on the feature information obtained from the projection model 230, the electronic device can obtain the feature information F p' can be obtained or calculated.
[0074]
number
[0075] Feature information F for numbers 7 and 8 p , F p' From the above, the electronic device obtains the characteristic information F p'' can be obtained or calculated.
[0076]
number
[0077] Number 9 is the F of number 8 p' For example, the number 9 corresponds to the self-attention of f c Using layer-based projection and linearization operations (LN), the feature information Fp' It can be defined to handle the addition operation of numbers 8 and 9 (e.g., +F p Calculation and / or +F' p The operation) can represent a residual connection (or identity mapping).
[0078] According to one embodiment, the electronic device can process the intrinsic information obtained from the sub-model 220 using the projection model 230. Within the projection model 230, the NCAP projector 410, the multi-head self-attention model 411, the first layer normalization model 412, the feedforward model 413, and the second layer normalization model 414 can be connected in a chain. Using the projection model 230, the electronic device can generate or obtain feature information F_p∈R̂(l×proj) from the intrinsic information. The feature information generated by the projection model 230 can include non-categorical information recognized by the sub-model 220 from the input image 402.
[0079] According to an embodiment, the electronic device may perform multi-head cross-attention between feature information [F_v∈R]^(c×hw) of the shallow CNN 421 and feature information F_p∈R^(l×proj) of the projection model 230 in the multi-head cross-attention model 423 of the image restoration model. Fp''' in Equation 10 may represent the feature information output from the multi-head cross-attention model 423.
[0080]
number
[0081] The query for performing multi-head cross-attention in Equation 10 may correspond to the feature information R^(c×hw) of the shallow CNN 421. d in Equation 10 may represent the dimension of the key vector. Qv in Equation 10 is a projection of Fv in Equation 5 (e.g., a projection based on the fc layer) and may represent a query vector. K_p^'' and V_p^'' are projections of Fp" in Equation 9 (e.g., a projection based on the fc layer) and may represent a key vector and a value vector, respectively. LN in Equation 10 may represent a linearization operation. The Qv[K_p^'']^T operation in Equation 10 may represent an attention score of self-attention. The T operation in Equation 10 may represent a matrix transpose operation.
[0082] The key and value for performing multi-head cross-attention in Equation 10 may have a size of R^(l×c) (e.g., [K_p^''∈R]^(l×c)). R^(l×c) may represent the feature dimension (number) of the shallow CNN 421. Q_v[K_p^'']^T in Equation 10 is = R^(hw×l), and [Q_v·K_p^'']^T·V_p^'' in Equation 10 may be R^(hw×c). Referring to Equation 10, feature information Fp''' obtained using a softmax operation and an LN (layer normalization) operation can be obtained from the multi-head cross-attention model 423.
[0083] For the feature information Fp''' obtained from the multi-head cross-attention model 423, the electronic device may perform a calculation represented by a chain connection of a merge model 424, a first layer normalization model 425, a feedforward model 426, and a second layer normalization model 427. Referring to FIG. 4, a residual connection for element-wise sum may be formed between the first layer normalization model 425 and the second layer normalization model 427. The residual connection may be formed between the first layer normalization model 425 and the second layer normalization model 427 independently of the feedforward model 426.
[0084] 4 , for the information obtained from the second layer normalization model 427, the electronic device may perform calculations based on the BiLSTM model 430 N times (e.g., 5 times). The combination of the first convolutional model 428, the second convolutional model 429, and the BiLSTM model 430 connected to the second layer normalization model 427 may be referred to as a decoder 470. In one embodiment, the feature information F obtained from the second layer normalization model 427 and input to the decoder 470 may be expressed as follows:
[0085]
number
[0086] W in Number 11 fmay represent a layer defined for a projection operation and the operation of that layer as an fc layer (or the weight of the fc layer). The decoder 470 may have a structure (sequential-recurrent block, SRB) in which the calculation represented by the BiLSTM model 430 is repeatedly executed N times. The electronic device 101 may use the pixel shuffle model 431 to increase the resolution and / or size of an image output by the decoder 470 (e.g., an image represented by the feature information F in Equation 11). For example, the output image 403 output from the pixel shuffle model 431 of the image restoration model may be determined based on Equation 12.
[0087]
number
[0088] When training an image restoration model having the structure of FIG. 4 (e.g., the second step of the training process), a loss function used to train the image restoration model may indicate the difference between the truth image corresponding to the input image 402 and the output image 403. For example, the L1 distance (e.g., Manhattan distance and / or rectangular street grid) between the truth image and the output image 403 may be determined as the loss function. Embodiments are not limited thereto, and L2 distance (or mean squared loss), SSIM (structural similarity index), TSSIM (triple SSIM), and KL (Kullback-Leibler) divergence loss functions for knowledge distillation may be used. For example, a loss function L_s based on the L2 distance may be defined as follows:
[0089]
number
[0090] I in the number 13 SR can represent the output image 403, and I HR can represent the truth image. For training an image restoration model based on structural information of text, a loss function based on TSSIM can be used, such as the loss function L_tssim in Equation 14.
[0091]
number
[0092] In Equation (14), x may correspond to the degraded output image 403, y may correspond to the output image 403, and z may correspond to the true image. μ and σ in Equation (14) are the mean and standard deviation of the corresponding images (e.g., x, y, z), respectively. C in Equation (14) may be an epsilon value (e.g., a specified number set to prevent zero division error).
[0093] According to one embodiment, the electronic device may use the pre-trained submodel 220 to train an image restoration model. The image restoration model may include an encoder including a TPS 420 and a shallow CNN 421 for extracting feature information from an input image 402. The image restoration model may include a synthesis layer (e.g., a multi-head cross-attention model 423) for combining the feature information with intrinsic information from an intermediate layer before the output layer of the submodel 220 that receives the input image 402. The image restoration model may include a decoder (e.g., a combination of a first convolutional model 428, a second convolutional model 429, and a BiLSTM model 430) connected to the synthesis layer for generating an output image 403 having a second resolution greater than the first resolution of the input image 402. The trained image restoration model may be provided as part of a software application for image restoration (e.g., the image restoration program 125 of FIG. 1 ).
[0094] In the following, an exemplary structure of an image restoration model connected to the teacher model 220 of FIG. 2 and / or FIG. 3 will be described with reference to FIG.
[0095] 5 shows an example block diagram of an image restoration model connected with a teacher model 210. The electronic device 101 and / or the processor 110 of FIG. 1 can execute the image restoration program 125 to obtain, generate, and / or train the image restoration model described with reference to FIG. 5.
[0096] As described above with reference to Figure 3, the output data of the submodel 220 may include a projection of the output data of an RNN decoder (e.g., the RNN decoder 350 in Figure 3) as shown in Equation 2. The input of the NCAP projector 410 may include the entire intermediate state vector (e.g., hidden state vector) of the RNN decoder at each of the multiple timings. If the submodel 220 performs parallel decoding, the input of the NCAP projector 410 may include the entire feature information obtained by the parallel decoding.
[0097] The output data of the teacher model 210 that receives the image 501 can be expressed as in Equation 15.
[0098]
number
[0099] The output data of the sub-model 220 may have the relationship of Equation 16. HR can represent the output of the teacher model 210 when a high-resolution image is input. For example, t HR is the information processed sequentially by the encoder and decoder of the STR, and can represent the information projected by the fc layer (e.g., the probability distribution of the text). For example, W in Equation 15 c is f c The layer can be represented as x LR can represent a low-resolution image.
[0100]
number
[0101] t in number 16 LR may represent the output of sub-model 220 when a low-resolution image is input. For example, number 16 may represent the output data of sub-model 220 that receives image 502. Based on the intrinsic information obtained from the submodel 220, the electronic device calculates p NCAPcan be obtained from the NCAP projector 410.
[0102] When training the sub-model 220, the electronic device can utilize the loss function L_distill in equation (17) to reduce the domain gap of the sub-model 220's prior knowledge (e.g., the domain difference between the high-resolution output image 503 and the low-resolution input image 502).
[0103]
number
[0104] Equation (2) of Equation (17) may be a loss function based on a smooth (e.g., kl divergence smoothness) profile based on temperature scaling, and equation (1) of Equation (17) may be a loss function based on a sharp (e.g., kl divergence smoothness) profile based on |t_HR-t_LR|_1. The electronic device may determine the loss function L_distill using either of the two equations of Equation (17). tHR and tLR of Equation (17) may represent prior knowledge obtained by inputting high-resolution and low-resolution images, respectively, into an STR including the submodel 220. tHR of Equation (17) may be generated from a frozen teacher model 210. tLR of Equation (17) may be generated from a submodel 220 in a trainable state. The loss function of Equation (17) may be determined by other methods (e.g., L1 distance). Referring to Equation 17, the truth data t_HR(τ) of the soft labels can be used to determine the loss function L_distill. HR A smoother profile can be created using [beta] and [tau] in Equation 17, which are parameters for adjusting the soft label and smoothness, can be set to, for example, 0.7 and 5, respectively. The loss function L_str of the submodel 220 can be determined by the loss function L_aux of Equation 18, which maximizes the cross-entropy loss (e.g., the cross-entropy of Equation 3) and margin with the truth image, as in Equation 18. For example, to use hard labels for training, one of the two formulas of Equation 18 can be used.
[0105]
number
[0106] When equation (1) of Equation (18) is used, the cross-entropy loss L_CE used for scene text recognition can be used. When equation (2) of Equation (18) is used, an additional loss function L_aux that maximizes the second margin can be used. gt may represent the ground truth label (e.g., ground truth data) of the image received as input. In the exemplary embodiment of FIG. 5, y gt can be "recycled".
[0107] The loss function L_aux in Equation 18 may correspond to L_aux in Equation 4. To the sub-model 220 trained by the loss function L_str in Equation 18, the electronic device may further apply a loss function that reduces the difference in attention scores between the high-resolution image and the low-resolution image and / or a loss function that reduces the difference between the probability distribution of the high-resolution image and the probability distribution of the low-resolution image. For example, the electronic device may utilize a loss function to focus on regions associated with one or more characters within the input image 502. For example, using weighted cross entropy (WCE) as in Equation 19, the electronic device may increase the weight of character strings that may be confused.
[0108]
number
[0109] For example, α in Equation 19 can be set to a value of 10, and β can be set to a value of 0.0005. ||A_HR-A_SR ||_1 in Equation 19 can mean the L1 distance. HR , A SR can represent the attention information (e.g., attention map) of the high-resolution image and the low-resolution image, respectively. pred can represent the output image 503. The loss function L txt is the attention information A of the low-resolution image. SR and attention information A of high-resolution images. HR can be defined to reduce the difference between
[0110] In one embodiment, the electronic device can at least partially train the image restoration model using a combination of the aforementioned loss functions (e.g., joint learning). The combination of loss functions, L_total, may be set as follows: Using L_total in Equation (20), which is an example of a WCE, backpropagation can be performed on the entire model starting from the pixel shuffle model 431. This backpropagation can be performed to reduce errors in the attention map and logits. For example, the loss function L_txt in Equation (19) may be used to reduce errors between the attention map and text logit information obtained from an additional artificial intelligence model (e.g., a text recognition network) for processing the image 502. The backpropagation can train the entire image restoration model to reduce the difference between the image 403 and the image 502.
[0111]
number
[0112] The parameters can be set to values such as λ1=1, λ2=1, λ3=0.01, and α=0.5 in Equation 20. α in Equation 20 can be a parameter for adjusting the training ratio between L_distill in Equation 17 and L_str in Equation 18. The embodiment is not limited thereto. L_s in Equation 20 may be determined as in Equation 13. L_tssim in Equation 20 may be defined as in Equation 10. L_distill in Equation 20 may be defined as in Equation 17. L_str in Equation 20 may be defined as in Equation 18. L_txt in Equation 20 may be defined as in Equation 19.
[0113] According to one embodiment, the electronic device can execute an image restoration model including a model 510 for restoring an input image 502, a sub-model 220, and a projection model 230 that can be executed at least temporarily concurrently with the model 510. The model 510 can be combined with any sub-model 220 for pre-trained character recognition. Using the sub-model 220, the electronic device can effectively obtain prior knowledge (or prior information) that can be used to restore or enhance the input image 502.
[0114] In the following, the performance of an image restoration model configured to obtain an output image 503 from an input image 502 will be described with reference to FIGS.
[0115] 6 shows graphs 611, 612, 621, and 622 illustrating the performance of an electronic device implementing an image restoration model, according to one embodiment. The graphs in FIG. 6 may be empirical graphs illustrating the performance of the image restoration model implemented by the electronic device and / or image restoration program of FIGS. 1-5.
[0116] Referring to FIG. 6, graph 611 shows the average number of top 5 predictions for characters (e.g., numbers 0-9 and alphabets a-z), and graph 612 shows the number of top 1 predictions. Referring to FIG. 6, graph 621 shows the standard deviation of the top 5 predictions for characters, and graph 622 shows the standard deviation of the top 1 predictions. An increase in standard deviation means that there are fewer values near the average, and therefore may indicate an increase in the accuracy of recognizing characters from an image. For example, Table 1 may show the average and standard deviation of the character prediction results obtained by running a submodel (e.g., submodel 220 of FIG. 2) on all characters.
[0117] [Table 1]
[0118] "Ours" in Table 1 may indicate the mean and standard deviation of the predicted results obtained by running an image restoration model according to one embodiment, and "Baseline" in Table 1 may indicate the mean and standard deviation of the predicted results obtained by running another model different from the image restoration model according to one embodiment.
[0119] FIG. 7 illustrates graphs 711, 712, 721, 722, 731, 732, 741, and 742 for illustrating the performance of an electronic device implementing an image restoration model according to one embodiment. Referring to FIG. 7 , regions 710, 720, 730, and 740 are illustrated, each corresponding to a specified character (e.g., 5, 9, u, and k) from a plurality of images. For example, region 710 illustrates graph 712 showing the average of the top five predictions for the specified character "5" and graph 711 showing the baseline. For example, region 720 illustrates graph 722 showing the average of the top five predictions for the specified character "9" and graph 721 showing the baseline. For example, region 730 illustrates graph 732 showing the average of the top five predictions for the specified character "u" and graph 731 showing the baseline. For example, region 740 illustrates graph 741 showing the average of the top five predictions for the specified character "k" and graph 742 showing the baseline. Referring to regions 710, 720, 730, and 740, within the results of predicting each character, the prediction value of the most predicted character can have a high value, and the frequency of predicting (e.g., confusing) characters below the Top 2 can be reduced.
[0120] For example, Table 2 can include the mean and standard deviation of the results of predicting the specified character using the sub-models.
[0121] [Table 2]
[0122] For example, Table 3 may include the standard deviation of the results of predicting the specified character using the sub-models.
[0123] [Table 3]
[0124] To determine whether the prior knowledge generated by the submodel is biased, the electronic device can calculate the relationship between the prior knowledge accuracy and the STISR accuracy, which can be expressed as the Pearson Correlation Coefficient (21).
[0125]
number
[0126] X in Equation 21 may represent the WER (word error rate) of the output and / or logits of the submodel. X in Equation 21 may be defined as the WER and CER of the text logits of the student recognizer. Y in Equation 21 may be defined as Y=STISR WER and CER. CER may be a character error rate (e.g., character error rate). n in Equation 21 may represent the number of total data, and i may represent an index defined to perform the summation operation.
[0127] In one embodiment, Table 4 can show the Pearson relationship between prior knowledge and STISR accuracy.
[0128] [Table 4]
[0129] Referring to Table 4, when compared with an existing method (baseline), the Pearson correlation coefficient of an image restoration model (e.g., Ours) according to one embodiment can be relatively reduced in both WER and CER. The reduction in the Pearson correlation coefficient of the image restoration model may mean that the image restoration model does not depend on incomplete information (e.g., prior knowledge).
[0130] In one embodiment, Table 5 can show the relationship between performance improvements of electronic devices and parameter increments.
[0131] [Table 5]
[0132] Referring to Table 5, performance improvement can be expected when the parameters of the image restoration model are increased by about 0.3%. According to one embodiment, the electronic device can use a commonly used adapter (e.g., a multi-layer perceptron (MLP)) and / or a convolutional adapter to implement the image restoration model. In one embodiment, performance can be relatively improved when a 1x1 convolutional adapter is used.
[0133] As described above, according to one embodiment, an electronic device can execute an image restoration model configured to generate text-related information (e.g., a text probability map) from an image. The image restoration model can include sub-models pre-trained to generate the information from the image. The image restoration model can restore or augment the image with intrinsic information used to generate extrinsic information output from the sub-models (e.g., one or more characters associated with the image and the relative positions of the one or more characters). By using text-related information to restore an image, the electronic device can be trained to interpret license plates and / or cover plates.
[0134] In the following, with reference to Fig. 8a and / or Fig. 8b, a license plate reconstructed by the image reconstruction model is exemplarily shown.
[0135] 8a and / or 8b show at least one license plate (license plate or number plate) as a subject included in an image reconstructed by an image reconstruction model, according to one embodiment.
[0136] 8a, an image 810 including at least one license plate obtained from an image restoration model is shown. Image 810 may be output or provided from an electronic device that has executed the image restoration model as a result of restoring or enhancing a low-resolution input image (e.g., input image 202 of FIG. 2).
[0137] For example, the electronic device may generate image 820 including a license plate based on Korean law. Image 820 may include a number representing the type of vehicle (e.g., 12), an alphabet representing the vehicle's use (e.g., "ka"), and a number representing a serial number uniquely assigned to the vehicle (e.g., 1234). For example, the electronic device may obtain image 830 including a license plate based on Korean law. Image 830 may further include characters representing a region associated with the license plate (e.g., a place name such as "Seoul") relative to image 820. The background color of the license plate represented by images 820, 830 may represent a vehicle category (e.g., private vehicle) defined by Korean law.
[0138] For example, the electronic device may generate image 840 including a license plate according to Chinese law. Image 840 may include a character representing the region associated with the license plate (e.g., Jing), a character representing the city associated with the license plate (e.g., a subregion of that region) (e.g., N), or other information related to the region or use. Image 840 may include a serial number uniquely assigned to the vehicle (e.g., 888R8). The color of the license plate represented by image 840 may represent the vehicle category (e.g., passenger car, large vehicle, bus, truck, and / or motorcycle).
[0139] For example, the electronic device may generate image 850 including a license plate under European Union law. Image 850 may include a symbol representing the European Union, letters indicating the region associated with the license plate (e.g., EST), and a serial number uniquely assigned to the vehicle to which the license plate is attached (e.g., "307RTB"). Embodiments are not limited thereto, and image 850 may further include the flag of a country in which the vehicle to which the license plate is attached is registered as a member of the European Union.
[0140] For example, the electronic device may generate an image 860 that includes a license plate under Japanese law. The image 860 may include characters representing a region (e.g., Tama), a number representing a vehicle category (e.g., 500), characters indicating the business use associated with the vehicle, and a serial number (e.g., 46-49) uniquely assigned to the vehicle to which the license plate is attached.
[0141] Referring to FIG. 8b, an image 870 including a U.S. legal license plate generated by an electronic device is shown, according to one embodiment. Referring to image 870, license plates can be generated based on U.S. law and include images and / or graphics defined by U.S. state governments. The license plate can include text representing the state government (e.g., "TEXAS," "ALABAMA," "KENTUCKY," etc.) along with an image and / or graphics representing the state government in which the vehicle is registered. Along with the text, the image representing the license plate can include a serial number uniquely assigned to the vehicle (e.g., a combination of letters and / or numbers, such as "GV71P").
[0142] FIG. 9 is a diagram illustrating the overconfidence phenomenon. When training the super-resolution networks 241, 242, 243, and 244 of FIG. 2 and the character recognition network (e.g., sub-model 220 of FIG. 2) simultaneously, the overconfidence phenomenon may occur due to a difference in training speed between the super-resolution network and the character recognition network. The overconfidence phenomenon may include predicting an incorrect character with a high probability from an image containing a character that is difficult to infer. The overconfidence phenomenon may adversely affect the results (e.g., character probability distribution) of the character recognition network (e.g., sub-model 220 of FIG. 2). According to one embodiment, the electronic device may reduce the overconfidence phenomenon using a loss function that combines hard-level truth data and soft labels (e.g., the output of a teacher model), as described above with reference to Equations (17) and (18).
[0143] 9, graphs 910 and 920 are shown representing confidence at the word level and character level, respectively. The x-axis of graphs 910 and 920 may represent a probability value corresponding to a character recognized by the neural network. The y-axis of graphs 910 and 920 may represent the probability (e.g., accuracy) that the character recognized by the neural network is correct. As shown in baselines 919 and 929 of graphs 910 and 920, the more proportional the probability value and accuracy, the less the overconfidence phenomenon.
[0144] Line 911 of graph 910 may represent an ideal relationship between the accuracy and reliability of an image restoration model of an electronic device. Line 912 of graph 910 may represent the accuracy of an image restoration model trained based on soft labels. Line 913 of graph 910 may represent the accuracy of an image restoration model trained based on hard labels. Line 921 of graph 920 may represent an ideal relationship between the accuracy and reliability of an image restoration model of an electronic device. Line 922 of graph 920 may represent the accuracy of an image restoration model trained based on soft labels. Line 923 of graph 920 may represent the accuracy of an image restoration model trained based on hard labels. Training with only hard labels may result in reduced accuracy compared to probability values. Training with only soft labels may reduce the overconfidence phenomenon but may result in reduced performance. Referring to graphs 910, 920 in Figure 9, in one embodiment, lines 911, 921 showing the accuracy of the image restoration model of the electronic device can be positioned closer to baselines 919, 929 showing more ideal accuracy than other lines 912, 913, 922, 923 (e.g., baselines showing the accuracy of the image restoration model in which the overconfidence phenomenon is minimized).
[0145] In one embodiment, a model trained to output extrinsic information, such as a text probability map, may be used to seek to increase or enhance the resolution of an image in which one or more characters have been captured. In one embodiment, intrinsic information from an intermediate layer within a model trained to output extrinsic information may be used to seek to increase or enhance the resolution of an image in which one or more characters have been captured. As noted above, in accordance with one embodiment, a non-transitory computer-readable storage medium may be provided that includes instructions. The instructions, when executed individually or collectively by at least one processor of an electronic device, may cause the electronic device to receive a request to restore a first image at a first resolution to an image at a second resolution that exceeds the first resolution. The instructions, when individually or collectively executed by the at least one processor, can cause the electronic device to execute, based on the received request, an image restoration model including an encoder for extracting feature information from the first image, a sub-model for determining a text probability map for the first image, a synthesis layer positioned before an output layer trained to output the text probability map for combining the feature information with intrinsic information in an intermediate layer of the sub-model, and a decoder connected to the synthesis layer for generating the image at the second resolution. When the instructions, when individually or collectively executed by the at least one processor, can cause the electronic device to provide, in response to the request, a second image at the second resolution obtained based on execution of the image restoration model. According to one embodiment, the electronic device can increase or enhance the resolution of an image in which one or more characters are captured using a model trained to output extrinsic information, such as a text probability map. According to one embodiment, the electronic device can increase or enhance the resolution of an image in which one or more characters are captured using intrinsic information in an intermediate layer in a model trained to output extrinsic information.
[0146] For example, the instructions, when executed individually or collectively by the at least one processor of the electronic device, can cause the electronic device to execute the image restoration model including the synthesis layer connected to the intermediate layer to extract the intrinsic information used to determine the text probability map, which is extrinsic information.
[0147] For example, the sub-model may be trained to output the text probability map representing one or more characters that appear to be captured by the first image and the location of the one or more characters.
[0148] For example, the sub-models can be pre-trained by a teacher model that is run with more parameters than the parameters for the sub-models.
[0149] For example, the instructions, when individually or collectively executed by the at least one processor of the electronic device, can cause the electronic device to receive a first signal including the request and a third image from an external electronic device via a communication circuit of the electronic device. The instructions, when individually or collectively executed by the at least one processor of the electronic device, can cause the electronic device to segment, based on receiving the first signal, a portion of the third image associated with a license plate as the first image. The instructions, when individually or collectively executed by the at least one processor of the electronic device, can cause the electronic device to transmit a second signal including the second image to the external electronic device based on obtaining the second image from the reconstruction model performed using the segmented first image.
[0150] As described above, according to one embodiment, an electronic device may include a memory storing instructions and at least one processor configured to execute the instructions. The instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to receive a request to restore a first image at a first resolution to an image at a second resolution greater than the first resolution. The instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to execute, based on the received request, an image restoration model including: an encoder for extracting feature information from the first image; a sub-model for determining a text probability map for the first image; a fusion layer located before an output layer trained to output the text probability map for combining implicit information of an intermediate layer of the sub-model with the feature information; and a decoder connected to the fusion layer for generating the image at the second resolution. The instructions, when executed individually or collectively by the at least one processor, can cause the electronic device to provide, in response to the request, a second image at the second resolution obtained based on execution of the image restoration model.
[0151] For example, the instructions, when executed individually or collectively by the at least one processor, can cause the electronic device to execute the image restoration model including the synthesis layer connected to the intermediate layer to extract the intrinsic information used to determine a text probability map, which is extrinsic information.
[0152] For example, the sub-model may be trained to output the text probability map representing one or more characters that appear to be captured by the first image and the location of the one or more characters.
[0153] For example, the sub-models can be pre-trained by a teacher model that is run with more parameters than the parameters for the sub-models.
[0154] For example, the instructions, when executed individually or collectively by the at least one processor, can cause the electronic device to receive a first signal from an external electronic device via a communication circuit of the electronic device, the first signal including the request and a third image. The instructions, when executed individually or collectively by the at least one processor, can cause the electronic device to segment, based on receiving the first signal, a portion of the third image associated with a license plate as the first image. The instructions, when executed individually or collectively by the at least one processor, can cause the electronic device to transmit a second signal including the second image to the external electronic device based on obtaining the second image from the reconstruction model performed using the segmented first image.
[0155] As described above, according to one embodiment, a method of an electronic device may be provided. The method may include an operation of obtaining a sub-model trained to receive an image and output a text probability map representing one or more characters associated with the image. The method may include an operation of using the sub-model to train an image restoration model including an encoder for extracting feature information from an input image, a synthesis layer receiving the input image and combining the feature information with intrinsic information in an intermediate layer preceding an output layer of the sub-model, and a decoder connected to the synthesis layer for generating an output image having a second resolution greater than a first resolution of the input image. The method may include an operation of providing the image restoration model as part of a software application for restoring images.
[0156] For example, the image restoration model may include the synthesis layer coupled to the hidden layer to extract the intrinsic information used to determine the text probability map, which is extrinsic information.
[0157] For example, the sub-model may be trained to output the text probability map representing one or more characters that appear to be captured by the input image and the location of the one or more characters.
[0158] For example, the obtaining operation may include obtaining the sub-model using a teacher model that is run using more parameters than parameters for the sub-model.
[0159] For example, the providing operation may include executing the image restoration model in response to a request to restore a portion associated with a license plate segmented from a source image.
[0160] As described above, according to one embodiment, an electronic device may include a memory storing instructions and at least one processor configured to execute the instructions. When the instructions, individually or collectively, are executed by the at least one processor, the electronic device may obtain a sub-model trained to output a text probability map representing one or more characters associated with an image based on receiving the image. When the instructions, individually or collectively, are executed by the at least one processor, the electronic device may perform training of an image restoration model using the sub-model, the image restoration model including an encoder for extracting feature information from an input image, a synthesis layer that receives the input image and combines the feature information with intrinsic information from an intermediate layer preceding an output layer of the sub-model, and a decoder connected to the synthesis layer for generating an output image having a second resolution greater than a first resolution of the input image. When the instructions, individually or collectively, are executed by the at least one processor, the electronic device may provide the image restoration model as part of a software application for restoring images.
[0161] For example, the image restoration model may include the synthesis layer connected to the hidden layer to extract the intrinsic information used to determine the text probability map, which is extrinsic information.
[0162] For example, the sub-model may be trained to output a text probability map representing one or more characters that appear to be captured by the input image and the location of the one or more characters.
[0163] For example, the instructions, when executed individually or collectively by the at least one processor, can cause the electronic device to obtain the sub-model using a teacher model that is run using more parameters than the parameters for the sub-model.
[0164] For example, the instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to execute the image restoration model in response to a request to restore a portion associated with a license plate segmented from a source image.
[0165] The devices described above may be implemented using hardware components, software components, and / or a combination of hardware and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. Furthermore, the processing device may access, store, manipulate, process, and generate data in response to the execution of software. For ease of understanding, a processing device may be described as using a single processing element, but those skilled in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, the processing device may include multiple processors or one processor and one controller. Additionally, other processing configurations are possible, such as parallel processors.
[0166] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure or instruct a processing device to operate in a desired manner, either individually or collectively. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device to be interpreted by or provide instructions or data to a processing device. The software may be distributed across computer systems connected to a network and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable storage media.
[0167] The method according to the embodiment may be implemented in the form of program instructions that can be executed by various computer means and recorded on a computer-readable medium. In this case, the medium may permanently store the computer-executable program or temporarily store it for execution or download. Furthermore, the medium may be various recording or storage means in the form of a single or multiple hardware devices combined together, and is not limited to media directly connected to any computer system, but may also be distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; ROMs, RAMs, flash memories, and the like, configured to store program instructions. Other examples of media include recording media or storage media managed by application stores that distribute applications, or by sites or servers that provide or distribute various other software.
[0168] Although the embodiments have been described above with reference to limited embodiments and drawings, those skilled in the art will appreciate that various modifications and variations may be made from the foregoing description. For example, the techniques described may be performed in an order different from that described, and / or the components of the described systems, structures, devices, circuits, etc. may be combined or combined in a manner different from that described, or may be substituted or replaced by other components or equivalents, and still achieve suitable results.
[0169] Accordingly, other aspects, embodiments, and equivalents of the claims are within the scope of the following claims.
Claims
1. A non-transitory computer-readable storage medium containing instructions that, when executed individually or collectively by at least one processor of an electronic device, cause the electronic device to: receiving a request to restore a first image at a first resolution to an image at a second resolution greater than the first resolution; Based on the request received: an encoder for extracting feature information from the first image; a sub-model for determining a text probability map for the first image; a fusion layer for combining implicit information of the intermediate layers of the sub-model with the feature information, the fusion layer being placed before an output layer trained to output the text probability map; and a decoder coupled to the compositing layer for generating an image of the second resolution; Run an image restoration model including and a non-transitory computer-readable storage medium that causes a second image at the second resolution obtained based on execution of the image restoration model to be provided in response to the request.
2. The instructions, when individually or collectively executed by the at least one processor of the electronic device, cause the electronic device to:
2. The non-transitory computer-readable storage medium of claim 1, wherein the non-transitory computer-readable storage medium causes the image restoration model including the synthesis layer connected to the intermediate layer to execute in order to extract the intrinsic information used to determine the text probability map, which is explicit information.
3. 2. The non-transitory computer-readable storage medium of claim 1, wherein the sub-model is trained to output a text probability map representing one or more characters that appear to be captured by the first image and the locations of the one or more characters.
4. 4. The non-transitory computer-readable storage medium of claim 3, wherein the sub-models are pre-trained by a teacher model that is run with more parameters than parameters for the sub-models.
5. The instructions, when individually or collectively executed by the at least one processor of the electronic device, cause the electronic device to: receiving a first signal from an external electronic device via a communication circuit of the electronic device, the first signal including the request and a third image; Segmenting a portion of the third image associated with a license plate as the first image based on receiving the first signal; 2. The non-transitory computer-readable storage medium of claim 1, further comprising: causing a second signal including the second image to be transmitted to the external electronic device based on obtaining the second image from the reconstruction model performed using the segmented first image.
6. An electronic device comprising: a memory for storing instructions; and at least one processor configured to execute the instructions; The instructions, when executed by the at least one processor individually or collectively, cause the electronic device to: receiving a request to restore a first image at a first resolution to an image at a second resolution greater than the first resolution; Based on the request received: an encoder for extracting feature information from the first image; a sub-model for determining a text probability map for the first image; a fusion layer for combining implicit information of the intermediate layers of the sub-model with the feature information, the fusion layer being placed before an output layer trained to output the text probability map; and a decoder coupled to the compositing layer for generating an image of the second resolution; Run an image restoration model including and causing the electronic device to provide, in response to the request, a second image at the second resolution obtained based on execution of the image restoration model.
7. The instructions, when executed by the at least one processor individually or collectively, cause the electronic device to: The electronic device of claim 6, wherein the image restoration model including the synthesis layer connected to the intermediate layer is executed to extract the intrinsic information used to determine the text probability map, which is explicit information.
8. 7. The electronic device of claim 6, wherein the sub-model is trained to output a text probability map representing one or more characters that appear to be captured by the first image and the locations of the one or more characters.
9. 9. The electronic device of claim 8, wherein the sub-models are pre-trained by a teacher model that is run with more parameters than the parameters for the sub-models.
10. The instructions, when executed by the at least one processor individually or collectively, cause the electronic device to: receiving a first signal including the request and a third image from an external electronic device via a communication circuit of the electronic device; Segmenting a portion of the third image associated with a license plate as the first image based on receiving the first signal; The electronic device of claim 6, further comprising: a second signal including the second image, the second signal being transmitted to the external electronic device based on obtaining the second image from the reconstruction model performed using the segmented first image.
11. 1. A method of an electronic device, comprising: based on receiving an image, obtaining a sub-model trained to output a text probability map representing one or more characters associated with the image; Using the submodel: an encoder for extracting feature information from the input image; a synthesis layer for combining the feature information with intrinsic information from an intermediate layer before the output layer of the sub-model, which receives the input image; and a decoder coupled to the compositing layer for generating an output image having a second resolution that exceeds the first resolution of the input image; an operation of performing training of an image restoration model, including: The method includes the act of providing the image restoration model as part of a software application for restoring images.
12. The image restoration model is 12. The method of claim 11, including the synthesis layer connected to the hidden layer to extract intrinsic information used to determine the text probability map, which is extrinsic information.
13. The sub-model is 12. The method of claim 11, wherein the method is trained to output the text probability map representing one or more characters that appear to be captured by the input image and the locations of the one or more characters.
14. The obtaining operation includes: The method of claim 11 , comprising obtaining the sub-model using a teacher model run with more parameters than parameters for the sub-model.
15. The providing operation comprises: The method of claim 11 , including the act of executing the image restoration model in response to a request to restore a portion associated with a license plate segmented from a source image.
16. An electronic device comprising: a memory for storing instructions; and at least one processor configured to execute the instructions; The instructions, when executed by the at least one processor individually or collectively, cause the electronic device to: obtaining a sub-model trained to output a text probability map representing one or more characters associated with the image based on receiving the image; Using the submodel: an encoder for extracting feature information from the input image; a synthesis layer for combining the feature information with intrinsic information from an intermediate layer before the output layer of the sub-model that receives the input image; and a decoder coupled to the compositing layer for generating an output image having a second resolution that exceeds the first resolution of the input image; Perform training of an image restoration model including and causing the electronic device to provide the image restoration model as part of a software application for restoring images.
17. The image restoration model is 17. The electronic device of claim 16, comprising the synthesis layer connected to the hidden layer to extract intrinsic information used to determine the text probability map, which is extrinsic information.
18. The sub-model is 17. The electronic device of claim 16, wherein the electronic device is trained to output a text probability map representing one or more characters that appear to be captured by the input image and the locations of the one or more characters.
19. The instructions, when executed by the at least one processor individually or collectively, cause the electronic device to:
17. The electronic device of claim 16, wherein the sub-model is derived using a teacher model run with more parameters than for the sub-model.
20. The instructions, when executed by the at least one processor individually or collectively, cause the electronic device to:
17. The electronic device of claim 16, wherein the electronic device causes the image restoration model to execute in response to a request to restore a portion associated with a license plate segmented from a source image.