Image processing system, text conversion device, decoding device, image processing method, program, and recording medium
The image processing system converts images to text for high compression and uses text-based models to restore original image quality, addressing limitations in existing compression technologies.
Patent Information
- Application Number
- PCT/JP2024/040823
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2024-11-18
- Publication Date
- 2025-09-04
AI Technical Summary
Existing image compression technologies are limited in achieving high compression rates across various image types and formats, and there is a need for methods that can decode compressed images effectively for their original use.
An image processing system that converts original images into text information, allowing for higher compression rates, and then generates decoded images based on this text information, using text generation and image generation models to restore the original image quality.
Enables higher compression rates and effective decoding of images, regardless of their type or format, ensuring the decoded images retain the quality and characteristics of the original images.
Smart Images

Figure JP2024040823_04092025_PF_FP_ABST
Abstract
Description
Image processing system, text conversion device, decoding device, image processing method, program, and recording medium
[0001] One embodiment of the present invention relates to an image processing system, an image processing method, a program, and a recording medium used for compressing an image. The present invention also relates to a text conversion device and a decoding device that constitute the image processing system.
[0002] Images as digital data can be compressed (encoded) using various compression formats such as JPEG (Joint Photographic Experts Group), zip, and run-length encoding. Recently, technologies for compressing images at higher compression rates have been developed (see, for example, Patent Document 1). Patent Document 1 discloses a method for combining lossy compression processing using the JPEG format with lossless compression processing using the JBIG (Joint Bi-level Image Experts Group) format, thereby enabling efficient compression processing at a high compression rate while minimizing degradation in image quality.
[0003] JP 2011-010023 A
[0004] On the other hand, image compression technology is required to achieve even higher compression rates than conventional technologies, including the technology described in Patent Document 1. Furthermore, existing image compression standards and compression methods may be limited to specific image types and formats to which they can be applied. Therefore, an image compression technology that is incompatible with the type and format of an original image before compression may not be usable for compressing that original image. Furthermore, it goes without saying that when using an original image after compression, the compressed image must be decoded in a state suitable for use of the original image.
[0005] One embodiment of the present invention has been made in consideration of the above circumstances, and aims to provide an image processing system, an image processing method, a program, and a recording medium that are highly versatile and capable of compressing an original image at a higher compression rate. Another embodiment of the present invention aims to provide a text conversion device and a decoding device that constitute the image processing system.
[0006] The above object is achieved by an image processing system described in [1] below: [1] An image processing system including a processor, wherein the processor executes an input reception process for receiving an input of an original image, a text conversion process for outputting text information relating to text expressing the original image based on the original image, an information reception process for receiving the text information, and a decoding process for generating a decoded image of the original image based on the text information.
[0007] Furthermore, according to one embodiment of the present invention, it is possible to provide a text conversion device as set forth in [2] below and a decoding device as set forth in [3] below. [2] A text conversion device constituting the image processing system as set forth in [1], having a processor included in the processor and executing input acceptance processing and text conversion processing. [3] A decoding device constituting the image processing system as set forth in [1], having a processor included in the processor and executing information acceptance processing and decoding processing.
[0008] In one embodiment of the present invention, an image processing system according to any one of [4] to
[18] below can be provided. [4] The image processing system according to [1], wherein the processor further executes a generation process to generate a reduced image of the original image, an extracted image of a partial region of the original image, or related information related to a partial region of the original image. [5] The image processing system according to [4], wherein, in the decoding process, the processor generates a decoded image based on text information and the reduced image, extracted image, or related information generated in the generation process. [6] The image processing system according to any one of [1], [4], or [5], wherein the processor further executes a storage process to store information required for executing the decoding process in association with the text information. [7] The image processing system according to [6], wherein, in the decoding process, the processor inputs text information into an image generation model constructed by learning to generate a decoded image, and the information required for executing the decoding process is information on the image generation model. [8] The image processing system according to [4] or [5], wherein the processor further executes a storage process to store the reduced image, extracted image, or related information generated in the generation process in association with the text information. [9] The image processing system according to any one of [1] or [4] to [8], wherein in the decoding process, the processor inputs text information into the image generation model constructed by learning to generate a decoded image, and the processor further performs a first update process to update the image generation model by re-learning based on the original image and the decoded image.
[10] The image processing system according to any one of [1] or [4] to [9], wherein in the text conversion process, the processor inputs the original image into the text generation model constructed by learning to output the text information, and the processor further performs a second update process to update the text generation model by re-learning based on the text information, the original image, and the decoded image.
[11] The image processing system according to [1] or [4] to
[10] , wherein in the text conversion process, the processor outputs text information based on the original image and information related to the acquisition of the original image.
[12] The image processing system according to
[11] , wherein the original image is a photographed image, and in the text conversion process, the processor outputs text information based on the original image and information related to the photographing of the original image.
[13] The image processing system according to any one of [1] or [4] to
[12] , wherein, when the original image is a photographed image, in the decoding process, the processor generates a decoded image based on text information and information related to the photographing of the original image.
[14] The image processing system according to any one of [1] or [4] to
[13] , wherein, when the original image is a photographed image, in the text conversion process, the processor outputs text information based on the original image and information related to the photographing of the original image, and in the decoding process, the processor generates a decoded image based on the text information and information related to the photographing of the original image.
[15] The image processing system according to any one of [1] or [4] to
[14] , wherein, when the processor receives input of an image group consisting of multiple images in the input receiving process, the processor performs the text conversion process in a first pattern in which text information representing an image included in the image group is output for each image, or in a second pattern in which, when outputting text information representing a second image acquired temporally later than a first image included in the image group, text information representing differences between the second image and the first image is output.
[16] The image processing system according to [4], [5] or [8], wherein, when the processor receives input of an image group including a plurality of images in the input receiving process, the processor switches the execution pattern of the text conversion process and the generation process between a third pattern in which the text conversion process and the generation process are executed, and a fourth pattern in which the text conversion process is executed but the generation process is not executed, depending on the images included in the image group.
[17] The image processing system according to
[15] or
[16] , wherein the image group is a video consisting of a plurality of frame images.
[18] The image processing system according to [1] or any of [4] to
[17] , wherein the processor further executes a correction process to correct text information based on a user operation.
[0009] The above-mentioned object is also achieved by an image processing method described in
[19] or
[20] below.
[19] An image processing method in which a processor performs the steps of accepting input of an original image, outputting text information related to text representing the original image based on the original image, accepting the text information, and generating a decoded image of the original image based on the text information.
[20] The image processing method described in
[19] , in which the processor further performs the step of generating a reduced image of the original image, an extracted image of a partial area of the original image, or related information related to a partial area of the original image.
[0010] A program according to one embodiment of the present invention is a program for causing a processor to execute each step included in the image processing method according to
[19] or
[20] . A recording medium according to one embodiment of the present invention is a processor-readable recording medium on which a program for causing a processor to execute each step included in the image processing method according to
[19] or
[20] is recorded.
[0011] According to one embodiment of the present invention, by outputting text information from an original image, it is possible to compress the original image at a higher compression rate regardless of the type of image. Furthermore, by generating a decoded image of the original image based on the output text information, it is possible to use the decoded image based on the original image even after the original image is compressed. According to one embodiment of the present invention, it is possible to provide an image processing system, an image processing method, a program, and a recording medium that achieve the above-mentioned effects, as well as a text conversion device and a decoding device that constitute the above-mentioned image processing system.
[0012] FIG. 1 is a diagram illustrating an example of use of an image processing system according to a first embodiment of the present invention. FIG. 2 is a diagram illustrating the configuration of an image processing system according to a first embodiment of the present invention. FIG. 3 is a diagram illustrating the hardware configuration of a text conversion device according to a first embodiment of the present invention. FIG. 4 is a diagram illustrating the hardware configuration of a decoding device according to a first embodiment of the present invention. FIG. 5 is a conceptual diagram of a text generation model. FIG. 6 is a conceptual diagram of an image generation model. FIG. 7 is a diagram illustrating an image processing flow according to a first embodiment of the present invention. FIG. 8 is an explanatory diagram illustrating association between text information and information of an image generation model. FIG. 9 is a diagram illustrating the content of image processing according to a second embodiment of the present invention. FIG. 10 is a diagram illustrating the image processing flow according to a second embodiment of the present invention. FIG. 11 is a diagram illustrating the content of image processing according to a third embodiment of the present invention. FIG. 12 is a diagram illustrating an example of an image processing flow according to a fourth embodiment of the present invention. FIG. 13 is a diagram illustrating another example of an image processing flow according to the fourth embodiment of the present invention. FIG. 14 is a diagram illustrating the content of image processing according to a fifth embodiment of the present invention. FIG. 15 is a diagram illustrating a case where text conversion processing is performed using a first pattern in a sixth embodiment of the present invention. FIG. 16 is a diagram illustrating a case where text conversion processing is performed using a second pattern in a sixth embodiment of the present invention. FIG. 17 is a diagram illustrating the content of image processing according to a seventh embodiment of the present invention.
[0013] Specific embodiments of the present invention will be described below. For ease of explanation, the following description may be given in terms of a GUI (Graphic User Interface). Furthermore, since the basic data processing technologies (communication / transmission technologies, data acquisition technologies, data recording technologies, data processing / analysis technologies, machine learning technologies, image processing technologies, image display technologies, visualization technologies, etc.) required to realize the present invention are well-known technologies, a description of these technologies will be omitted.
[0014] In this specification, the concept of "device" includes not only a single device that performs a specific function, but also a combination of multiple devices that exist independently and in a distributed manner but cooperate (link) to perform a specific function. In addition, in this specification, the concept of "system" includes a system that is composed of multiple devices connected in a state where they can communicate with each other, and a system that is composed of a single device.
[0015] In this specification, the term "user" refers to a user of the image processing system of the present invention. Specifically, a user is someone who uses information obtained by the functions of the image processing system of the present invention (more specifically, text information and decoded images, etc., as described below).
[0016] In addition, in this specification, the term "person" refers to an entity that performs a specific action, and includes individuals, groups, corporations such as companies, and organizations, as well as computers and devices that constitute artificial intelligence (AI). Artificial intelligence (AI) is a technology that realizes intelligent functions such as inference, prediction, and judgment using hardware and software resources.
[0017] Additionally, in this specification, machine learning algorithms may include neural networks, convolutional neural networks, recurrent neural networks, attention, transformers, variational autoencoders, generative adversarial networks, deep learning neural networks, Boltzmann machines, matrix factorization, factorization machines, m-way factorization machines, field-aware factorization machines, field-aware neural factorization machines, support vector machines, Bayesian networks, decision trees, random forests, and other machine learning algorithms.
[0018] First Embodiment of the Present Invention In a first embodiment of the present invention (hereinafter referred to as the first embodiment), an image processing system of the present invention compresses an image by encoding the image, and decodes the encoded image. Hereinafter, the image to be encoded and decoded will be referred to as an “original image Po.”
[0019] Here, unless otherwise specified, the term "image" in this specification refers to digital image data (hereinafter referred to as image data) that defines the gradation values of each of the multiple pixels that make up an image. Image data includes raw image data before compression and image data after compression. Compressed image data includes image data that has undergone lossy compression, such as JPEG, and image data that has undergone lossless compression, such as GIF (Graphics Interchange Format) or PNG (Portable Network Graphics). Images also include images captured with a camera or other imaging device (captured images), images obtained by scanning an existing photograph with a scanner or the like (scanned images), images drawn using drawing software, CG (Computer Graphics), and edited images (edited images). Images may be still images or moving images, and moving images may be either with or without sound. The following description will use the case of a captured image as an example. Unless otherwise specified, captured images are considered to be still images.
[0020] In the first embodiment, as shown in FIG. 1, the image processing system receives an input of an original image Po, encodes the original image Po (converts it into text), and outputs text information Tx.
[0021] The text information Tx is information about the text that expresses the original image Po, for example, information about the text that expresses the content of the original image Po. More specifically, as shown in Fig. 1 , the text information Tx is a caption that explains the content of the original image Po, and is words and sentences that express the subject and background included in the original image Po, the scene that appears in the original image Po, and the theme that can be sensed by looking at the original image Po, and the like, and the content can be understood by a human being.
[0022] The text information Tx may also include information regarding the location and date / time of the original image Po, as well as information regarding the camera orientation and tilt angle (elevation angle) relative to the horizontal when the original image Po was captured. Furthermore, the text information Tx may also include information obtained by recognizing, using known voice recognition technology, the voice sound collected simultaneously with the capture of the original image Po, i.e., voice transcription information. Note that images converted into binary characters (specifically, 0 and 1) using Base64 do not qualify as the text information Tx of the present invention. Furthermore, a text data file (text file) representing the contents of the original image Po and a data file obtained by zip-compressing that text file qualify as text information Tx.
[0023] In the first embodiment, the original image Po is compressed by converting the original image Po into text, encoding it, and storing the resulting text information Tx. This makes it possible to reduce the data size after compression (encoding) compared to when the original image Po is compressed using a conventional compression method such as JPEG, zip, or run-length encoding and then stored as image data.
[0024] In the first embodiment, the image processing system generates a decoded image Pd of an original image Po based on text information Tx, as shown in FIG. 1 . The decoded image Pd is an image generated based on the text information Tx for the purpose of decoding the original image Po. Specifically, the decoded image Pd is an image that is identical to or similar to the original image Po, and is an image that can give the same impression and atmosphere as the original image Po. Here, decoding means returning to a state that is equivalent to or similar to the original state, and restoration and reproduction are synonymous with decoding. In other words, the decoded image Pd of the original image Po can also be said to be a reproduced or restored image of the original image Po.
[0025] In the first embodiment, the encoded (text-converted) original image Po is decoded by generating a decoded image Pd based on text information Tx. Then, by using the obtained decoded image Pd, it is possible to use the original image Po before compression, or more precisely, an image that is the same as or similar to the original image Po.
[0026] [Configuration example of image processing system] Next, a configuration example of an image processing system according to the first embodiment (hereinafter referred to as image processing system 10) will be described with reference to Figures 2 to 4. The image processing system 10 is made up of one or more computers, and for example, as shown in Figure 2, it is made up of a plurality of computers (two in the example shown in Figure 2) connected so as to be able to communicate via a network N. Note that the number of computers making up the image processing system 10 can be set to any number.
[0027] The computers that make up the image processing system 10 are computers that can be used by users, for example, client terminals, and are specifically configured as PCs (Personal Computers), smartphones, tablet terminals, wearable terminals, PDAs (Personal Digital Assistants), etc. Note that the computers that make up the image processing system 10 are not limited to terminals owned by users, and may be configured as terminals that are not owned by users but can be used by entering a PIN number, password, etc., or making a deposit, etc., when visiting a store, such as a store-installed terminal.
[0028] The computers constituting the image processing system 10 may also include a server computer, more specifically, a server computer for a cloud service. The server computer may be a server computer for an ASP (Application Service Provider), SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service). In this case, when necessary information is input into a client terminal, the server computer performs various processes and calculations based on the input information, and the calculation results are output on the client terminal. In other words, the functions provided by the server computer can be used on the client terminal.
[0029] 3 and 4, each computer constituting the image processing system 10 includes a processor 21, a memory 22, a communication interface 23, a storage 24, an input device 25, and an output device 26. In other words, the image processing system 10 includes at least one processor 21 and at least one memory 22.
[0030] The processor 21 is configured by, for example, a central processing unit (CPU), a micro-processing unit (MPU), a microcontroller unit (MCU), a graphics processing unit (GPU), a digital signal processor (DSP), a tensor processing unit (TPU), an application specific integrated circuit (ASIC), etc. The memory 22 is configured by, for example, semiconductor memory such as a read only memory (ROM) and a random access memory (RAM).
[0031] The communication interface 23 may be configured by, for example, a network interface card, a communication interface board, etc. Each computer constituting the image processing system 10 can communicate with other devices connected to a network N such as the Internet or a mobile communication line via the communication interface 23.
[0032] The storage 24 may be configured, for example, by a flash memory, a hard disc drive (HDD), a solid state drive (SSD), a flexible disc (FD), a magneto-optical disc (MO disc), a compact disc (CD), a digital versatile disc (DVD), a secure digital card (SD card), or a universal serial bus memory (USB memory). The storage 24 may be built into a computer constituting the image processing system 10 or may be externally attached to the computer. Alternatively, the storage 24 may be configured by a network attached storage (NAS) or the like. The storage 24 may also be an external device, such as an online storage or a database server, that can communicate with each computer constituting the image processing system 10 via the network N.
[0033] The input device 25 is a device that accepts input operations from the user and is configured, for example, by a touch panel, a keyboard, a mouse, etc. The input device 25 also includes a photographing device (image sensor) such as a built-in camera of a smartphone, a microphone for collecting sound, etc. The output device 26 is configured, for example, by a display, a speaker, etc.
[0034] Furthermore, software such as an operating system (OS) program and an application program for executing image processing is installed in each computer constituting the image processing system 10. These programs are read and executed by the processor 21, causing each computer constituting the image processing system 10 to perform its functions, specifically, to execute a series of processes related to the conversion (compression) of the original image Po into text and decoding.
[0035] More specifically, of the multiple computers that make up the image processing system 10, one or more computers function as a text conversion device 12 included in the image processing system 10. A processor 21 included in the text conversion device 12 executes input reception processing and text conversion processing.
[0036] The input reception process is a process of receiving input of an original image Po from a user. In the first embodiment, the method of inputting the original image Po is not particularly limited, but includes, for example, inputting data of the original image Po captured by a photographing device such as a camera, inputting data obtained by scanning a photograph of the original image Po with a scanner or the like, and inputting data of the original image Po drawn by drawing software or CG. The original image Po may also be input by downloading data of the original image Po from an external device or a web server or the like via the network N.
[0037] The text conversion process is a process of outputting text information Tx about an input original image Po based on the input original image Po. In the text conversion process, the processor 21 uses a text generation model Mt to create text that expresses the original image Po and outputs text information Tx related to the text. When an image is input, the text generation model Mt analyzes the input image to identify the features of the image and generates text according to the features. The generated text is then output as text information Tx.
[0038] The text generation model Mt is a model installed in a known caption generation AI, and has an input layer Mti, multiple intermediate layers Mtm, and an output layer Mto, as shown in Figure 5. The text generation model Mt is constructed by machine learning (equivalent to learning) similar to that used to construct a caption generation AI, and more specifically, each parameter in the model is determined by machine learning. Note that a text-to-image algorithm can be used for the text generation model Mt, and specifically, for example, VLMs (Vision-Language Models) such as CLIP (Contrastive Language-Image Pre-training) can be used.
[0039] The text generation model Mt may also be stored in an external computer (e.g., a server computer). In this case, the processor 21 of the text generator 12 can access the computer via the network N and use the text generation model Mt stored in that computer.
[0040] Furthermore, if information relating to the location, date and time when the original image Po was taken, and the camera settings at the time of taking the image (i.e., the photographing conditions) is recorded as tag information in the original image Po, the text conversion process may output text information Tx relating to the tag information. Furthermore, if sound is collected by a microphone or the like near the location where the original image Po was taken at the time the original image Po was taken, the sound may be recognized and text information Tx relating to text obtained by converting the sound into text may be output.
[0041] The method for outputting the text information Tx is not particularly limited, but for example, a data file of the text information Tx may be generated and saved in a predetermined storage location, or the text indicated by the text information Tx may be displayed on a display or the like.
[0042] Furthermore, the text information Tx may include, as additional information, information indicating that the text information Tx is used to generate a decoded image Pd of the original image Po. For example, the file name of the text information Tx may include character string information indicating that the information is used to generate a decoded image Pd of the original image Po, or an extension specific to such information.
[0043] Of the multiple computers that make up the image processing system 10, one or more computers function as a decoding device 14 included in the image processing system 10. A processor 21 included in the decoding device 14 executes information reception processing, decoding processing, and storage processing.
[0044] The information reception process is a process of receiving the text information Tx output by the text conversion device 12. Specifically, in the information reception process, the processor 21 of the decoding device 14 receives the data of the text information Tx from the text conversion device 12 via the network N, thereby receiving the text information Tx.
[0045] The decoding process is a process of generating a decoded image Pd of the original image Po based on the received text information Tx. In the decoding process, the processor 21 generates a decoded image Pd from the text information Tx using an image generation model Mg. When the text information Tx is input, the image generation model Mg analyzes the text information Tx, recognizes the text (caption) indicated by the text information Tx, and generates an image according to the recognition result.
[0046] The image generation model Mg is a model installed in a known image generation AI, and as shown in FIG. 6 , has an input layer Mgi, multiple intermediate layers Mgm, and an output layer Mgo. The image generation model Mg is constructed by machine learning (equivalent to learning) similar to that used to construct an image generation AI. More specifically, parameters within the model are acquired by machine learning. Note that the image generation model Mg can utilize a text-to-image algorithm, such as a variational autoencoder (VAE) or a generative adversarial network (GAN). Examples of image generation models using VAE include DALL-E, Stable Diffusion, Midjourney, and Adobe Firefly. These models, or models developed from these, can be used as the image generation model Mg.
[0047] The image generation model Mg may also be stored in an external computer (e.g., a server computer). In this case, the processor 21 of the decoding device 14 can access the computer via the network N and use the image generation model Mg stored in that computer.
[0048] The saving process is a process of saving information necessary for executing the decoding process in association with the text information Tx. Here, the information necessary for executing the decoding process is information about the image generation model Mg, more specifically, information about the image generation model Mg and information about the image generation model Mg itself.
[0049] The information about the image generation model Mg includes, for example, parameters within the image generation model Mg, specifically, hyperparameters. Hyperparameters include the number of epochs, the threshold, the number of hidden layers Mgm, and the number of neurons per hidden layer Mgm. If the image generation model Mg is stored on a computer other than the decoding device 14, for example, an external server computer, the information about the image generation model Mg may also include uniform resource locator (URL) information indicating the storage location. The information about the image generation model Mg itself also includes information for identifying the image generation model Mg, specifically, the number of layers included in the image generation model Mg, information about the association between layers, and information about the weighting (coefficients) between layers.
[0050] The information of the image generation model Mg saved in the saving process (hereinafter referred to as model information IM) is specific to the text information Tx associated with that model information IM. Specifically, the model information IM (more specifically, the parameters in the image generation model Mg) is set, or more precisely tuned, so that the decoded image Pd generated by inputting the text information Tx into that image generation model Mg matches or resembles the original image Po from which the text information Tx was generated.
[0051] 8, the model information IM is stored in association with the text information Tx, so that an appropriate image generation model Mg can be used when subsequently generating a decoded image Pd based on the text information Tx. Furthermore, the decoded image Pd generated using the image generation model Mg corresponds to the original image Po that was the source of generation of the text information Tx input to the image generation model Mg.
[0052] The mode of storing the model information IM in association with the text information Tx may be a mode in which the model information IM is written into the data file of the text information Tx, that is, a mode in which the text information Tx and the model information IM are stored in a single data file. Alternatively, a data file for the model information IM may be generated separately from the text information Tx, and the path to that data file or identification information for that file (e.g., a file name, etc.) may be stored in the data file for the text information Tx. Alternatively, the correspondence between the data file for the model information IM and the data file for the text information Tx may be stored as separate data, for example, table data.
[0053] Furthermore, the entity that executes the saving process is not limited to the processor 21 of the decoding device 14, but may also be the processor 21 of the text conversion device 12. In this case, the processor 21 of the text conversion device 12 acquires the model information IM data from the decoding device 14, and saves the acquired model information IM in association with the text information Tx that is input to the image generation model Mg related to the model information IM.
[0054] [Example of Image Processing Method According to First Embodiment] Next, as an example of the operation of the image processing system 10 in the first embodiment, an image processing flow using the image processing system 10 will be described. The image processing flow described below uses the image processing method of the present invention. In other words, each step (process) in the image processing flow described below corresponds to a component of the image processing method of the present invention. Note that the flow below is merely an example, and some steps in the flow may be deleted, new steps may be added, or the execution order of two steps in the flow may be reversed, as long as it does not deviate from the spirit of the present invention.
[0055] 7 by one or more processors 21 included in the image processing system 10, more specifically, by the processors 21 of the text conversion device 12 and the decoding device 14. That is, at each step in the image processing flow, the processor 21 of the device corresponding to that step, either the text conversion device 12 or the decoding device 14, reads an application program for image processing and executes data processing related to that step.
[0056] Specifically, in the image processing flow according to the first embodiment, the processor 21 of the text conversion device 12 first executes an input acceptance process to accept input of an original image Po from a user (S001). In this process, for example, a certain user (hereinafter referred to as the first user) captures an image with a digital camera, saves the captured image in a PC constituting the text conversion device 12, and launches an image processing application program on the PC. The first user then performs a predetermined input operation on the PC to input the captured image as the original image Po. This causes the processor 21 of the text conversion device 12 to accept the input of the original image Po.
[0057] Thereafter, the processor 21 of the text conversion device 12 executes a text conversion process and outputs text information Tx relating to the text expressing the original image Po based on the input original image Po (S002). In this step, the original image Po is input to the text generation model Mt, and the text information Tx is output as the result of compressing the original image Po. Note that the output text information Tx may include, as additional information, information indicating that the text information Tx is data for image decoding.
[0058] Next, when the user inputs text information Tx to request image decoding, the processor 21 of the decoding device 14 executes an information acceptance process and accepts the input text information Tx (S003). In this process, for example, a second user different from the first user receives the text information Tx transmitted from the first user via the network N, saves it in a PC constituting the decoding device 14, and starts an application program for image processing on the PC. As a result, the processor 21 of the decoding device 14 accepts the input text information Tx and the image decoding request from the second user.
[0059] Thereafter, the processor 21 of the decoding device 14 executes a decoding process to generate a decoded image Pd of the original image Po based on the input text information Tx (S004). In this step, the text information Tx is input to the image generation model Mg, and a decoded image Pd is generated as a result of decoding based on the text information Tx. The generated decoded image Pd is displayed on a display serving as an output device 26 of the decoding device 14, and the second user can confirm the decoded image Pd displayed on the display.
[0060] The processor 21 of the decoding device 14 also executes a storage process (S005). In this process, the text information Tx input in step S003 and information about the image generation model Mg (model information IM) used in the decoding process based on the text information Tx are associated with each other and stored in a predetermined storage location. As a result, when a decoding process is subsequently performed based on the same text information Tx, a decoded image Pd can be generated using the image generation model Mg identified by the model information IM associated with the text information Tx. As a result, even if multiple model providers each provide image generation models Mg of different standards, the same decoded image Pd can be generated from the text information Tx by using the model information IM associated with the text information Tx, regardless of the image generation model Mg of any standard.
[0061] When the above series of steps are completed, the image processing flow according to the first embodiment ends.
[0062] <<Second Embodiment of the Present Invention>> In the first embodiment, when generating a decoded image Pd of an original image Po, the decoded image Pd is generated based only on the text information Tx. However, this is not limited to this, and for example, the decoded image Pd may be generated by using a reduced image of the original image Po together with the text information Tx. This case is referred to as the second embodiment of the present invention, and the second embodiment will be described below. Note that in the second embodiment, differences from the first embodiment will be mainly described, and descriptions of points in common with the first embodiment will be omitted.
[0063] In the second embodiment, at least one processor 21 included in the image processing system 10, specifically, for example, the processor 21 of the text conversion device 12, further executes a generation process in addition to the input reception process and the text conversion process.
[0064] The generation process is a process of generating a reduced image Ps of the original image Po. The reduced image Ps is an image that has been reduced to a level appropriate for use as auxiliary data when the decoding device 14 generates a decoded image Pd of the original image Po. For example, a thumbnail image of the original image Po can be used as the reduced image P.
[0065] The generated reduced image Ps of the original image Po is then stored in association with text information Tx generated based on the same original image Po, as shown in Fig. 9. The generated reduced image Ps is stored in association with the text information Tx. The manner in which the reduced image Ps is stored in association with the text information Tx is similar to the manner in which the model information IM is stored in association with the text information Tx.
[0066] 9, in the decoding process, the processor 21 of the decoding device 14 generates a decoded image Pd based on the text information Tx and the reduced image Ps generated in the generation process and associated with the text information Tx. By generating the decoded image Pd based on the reduced image Ps of the original image Po together with the text information Tx in this way, it is possible to obtain a decoded image Pd that reproduces the original image Po with higher accuracy than when the decoded image Pd is generated from only the text information Tx.
[0067] Next, an image processing flow according to the second embodiment will be described with reference to FIG. 10. As shown in FIG. 10, the image processing flow according to the second embodiment is the same as the image processing flow according to the first embodiment (S011 to S012, S014 to S016), except for the addition of step S013, which executes a generation process. In the generation process, as described above, the processor 21 of the text conversion device 12 generates a reduced image Ps of the original image Po and stores the generated reduced image Ps in association with text information Tx. Note that known image processing techniques can be used to generate the reduced image Ps.
[0068] In the second embodiment, when the processor 21 of the decoding device 14 executes the decoding process to generate a decoded image Pd of the original image Po, the processor 21 generates the decoded image Pd using the text information Tx and the reduced image Ps.
[0069] <<Third Embodiment of the Present Invention>> In the above-described embodiment, the entire original image Po is converted into text (encoded), and text information Tx representing the entire original image Po is output. Furthermore, in the above-described embodiment, the decoded image Pd of the original image Po, which is generated based on the text information Tx, is generated by decoding the entire original image Po from the text information Tx. However, this is not limited to this, and whether or not to convert into text (encode) each region of the original image Po may be determined. This case is referred to as the third embodiment of the present invention, and the third embodiment will be described below. Note that the third embodiment will mainly describe the differences from the first and second embodiments, and a description of the points in common with these embodiments will be omitted.
[0070] In the third embodiment, a known segmentation technique is applied to the original image Po to divide the original image Po into one or more regions, thereby dividing the original image Po into a background region and a subject region, and if the original image contains multiple subjects, dividing it into regions for each subject.
[0071] Furthermore, if the original image Po contains a characteristic region (hereinafter referred to as a characteristic region Af), the characteristic region Af does not need to be converted into text (encoded). This is because the characteristic region Af in the original image Po is required to be decoded with high reproduction accuracy when the original image Po is decoded, and therefore it is appropriate to leave the image as it is without converting it into text (encoding).
[0072] Here, the characteristic area Af is an area in the original image Po that contains a particularly distinctive and important subject. Specifically, the characteristic area Af corresponds to an area in the original image Po where a person's face is located, an area where a central person is located if there are multiple people in the original image Po, or an area where a subject that is highly related to that person (for example, a pet that the person keeps) is located.
[0073] For the reasons mentioned above, it is preferable to store the position and range of the characteristic region Af in the original image Po, as well as the gradation value (pixel value) of each pixel contained in the characteristic region Af. That is, it is preferable to extract the image of the characteristic region Af in the original image Po (more specifically, the image fragment located in the characteristic region Af) from the original image Po and store the image as is, as shown in Figure 11. In this case, it is preferable to store the extracted image Pe of the characteristic region Af as an image compressed using a well-known compression method such as JPEG.
[0074] As described above, in the third embodiment, the aforementioned generation process is executed, and in the generation process, an extracted image Pe of the characteristic region Af is generated instead of or in addition to the reduced image Ps of the original image Po. Furthermore, in the third embodiment, a saving process is executed, and in the saving process, the extracted image Pe of the characteristic region Af is saved in association with text information Tx generated based on the original image Po including the characteristic region Af. The manner in which the extracted image Pe is saved in association with the text information Tx is similar to the manner in which model information IM is saved in association with the text information Tx.
[0075] On the other hand, the non-characteristic regions An in the original image Po are subjected to the above-described text conversion process, and are compressed and encoded as text information Tx. Non-characteristic regions An are regions of the original image Po other than the characteristic regions Af, and more specifically, regions such as background images where slight differences between the original image Po and the decoded image Pd have little effect on the impression of the image.
[0076] When determining whether each region in the original image Po is to be subjected to text conversion processing, for example, known subject recognition technology and scene recognition technology may be used to identify the subject and scene of each region in the original image Po, and a priority order may be assigned to each region according to the identification results. The priority order may be set according to a predetermined rule, and a region with a priority higher than a predetermined rank may be set as a characteristic region Af, and this region may be excluded from the text conversion processing.
[0077] Furthermore, whether or not to subject each region in the original image Po to text processing may be determined based on the sharpness of that region. Sharpness is a value indicating the degree to which the camera was in focus at the time the original image Po was captured. The higher the sharpness of a region, the more likely it is to be an important region in the original image. Therefore, regions whose sharpness exceeds a reference value may be set as characteristic regions Af and not be subject to text processing, and a generation process may be performed to generate an extracted image Pe of the characteristic regions Af. Conversely, regions whose sharpness does not exceed the reference value may be likely to be regions of low priority in the original image Po. Therefore, they may be set as non-characteristic regions An and subject to text processing, and the image of the non-characteristic regions An may be converted to text (encoded) and compressed as text information Tx.
[0078] In the third embodiment, the processor 21 of the decoding device 14 generates a decoded image Pd of the original image Po based on an extracted image Pe of the characteristic region Af and text information Tx obtained for the non-characteristic region An during decoding processing. Specifically, as shown in Fig. 11 , the processor 21 of the decoding device 14 generates the decoded image Pd by pasting the extracted image Pe of the characteristic region Af onto an image generated by inputting the text information Tx into an image generation model Mg. At this time, the position where the extracted image Pe is pasted in the decoded image Pd is set to correspond to the position of the characteristic region Af in the original image Po.
[0079] [Modifications of the Third Embodiment] Several modifications are possible for the above-described third embodiment. Four modifications (hereinafter, a third embodiment, a third embodiment, a third embodiment, and a third embodiment) of the third embodiment will be described below.
[0080] (Embodiment 3A) In the above-described third embodiment, the characteristic regions Af in the original image Po are excluded from the text conversion process, but this is not limiting. In Embodiment 3A, which is a modification of the third embodiment, both the characteristic regions Af and non-characteristic regions An are subject to the text conversion process. Furthermore, in Embodiment 3A, the level of detail (accuracy) of the text conversion process differs between the characteristic regions Af and the non-characteristic regions An. Specifically, the amount of text information Tx obtained by converting the characteristic regions Af of the original image Po into text is greater than the amount of text information Tx obtained by converting the non-characteristic regions An into text, and more specifically, the text information Tx contains a greater number of words. Furthermore, in Embodiment 3A, when the processor 21 of the decoding device 14 executes the decoding process to generate a decoded image Pd of the original image Po, the processor 21 inputs the text information Tx obtained for the characteristic regions Af and the text information Tx obtained for the non-characteristic regions An into the image generation model Mg, respectively, to generate the decoded image Pd.
[0081] In the third embodiment, which is another variation of the third embodiment, the processor 21 of the text generator 12 executes a generation process in which the processor 21 analyzes the characteristic regions Af in the original image Po to generate related information related to the characteristic regions Af. The related information is information indicating text related to the subject placed in the characteristic regions Af. If the subject is a person, the related information would include the person's age, gender, and external characteristics.
[0082] As a specific means for generating the related information, for example, the feature area Af in the original image Po may be analyzed using an analysis method such as a known object recognition technique or pattern recognition technique, and information corresponding to the analysis results may be acquired as the related information. Alternatively, AI may be used to identify information corresponding to the feature amount of the image of the feature area Af, and the information may be acquired as the related information.
[0083] In addition, in the third embodiment, a saving process is executed in which related information about the characteristic region Af is saved in association with text information Tx generated based on the original image Po including the characteristic region Af. The manner in which the related information is saved in association with the text information Tx is similar to the manner in which the model information IM is saved in association with the text information Tx.
[0084] In the third B embodiment, when the processor 21 of the decoding device 14 executes the decoding process to generate a decoded image Pd of the original image Po, the processor 21 inputs the related information acquired about the feature area Af and the text information Tx obtained by converting (compressing) the original image Po into text into an image generation model Mg, thereby generating the decoded image Pd.
[0085] (Embodiment 3C) In embodiment 3C, which is another variation of embodiment 3, the processor 21 of the text conversion device 12 executes a generation process for a characteristic region Af in an original image Po to generate an auxiliary image of the characteristic region Af, and also executes a text conversion process to output text information Tx for the characteristic region Af. The auxiliary image of the characteristic region Af is obtained by reducing the extracted image Pe of the characteristic region Af to a level appropriate for use as auxiliary data for generating a decoded image Pd. In other words, it is an image of the range of the reduced image Ps of the original image Po that corresponds to the characteristic region Af.
[0086] In addition, in the third embodiment, a saving process is executed in which an auxiliary image of the characteristic region Af is saved in association with text information Tx generated based on the original image Po including the characteristic region Af. The manner in which the auxiliary image is saved in association with the text information Tx is similar to the manner in which the model information IM is saved in association with the text information Tx.
[0087] In the third embodiment, when the processor 21 of the decoding device 14 executes a decoding process to generate a decoded image Pd of the original image Po, the processor 21 inputs the text information Tx and auxiliary image acquired for the characteristic region Af and the text information Tx acquired for the non-characteristic region An into the image generation model Mg to generate the decoded image Pd. This process can further improve the reproduction accuracy of the original image Po in the generated decoded image Pd. Note that known image generation techniques, specifically known image generation AI such as stable diffusion, can be used as a technique for generating a new image using both an image and text.
[0088] (Embodiment 3D) In embodiment 3D, which is a further development of embodiment 3C described above, the processor 21 of the text conversion device 12 performs a generation process for both the characteristic regions Af and the non-characteristic regions An in the original image Po to generate auxiliary images, and also performs a text conversion process to output text information Tx.
[0089] Furthermore, in the third embodiment, a saving process is executed in which auxiliary images of the characteristic regions Af and non-characteristic regions An are saved in association with text information Tx generated based on an original image Po including these regions. Saving auxiliary images in association with text information Tx is similar to saving model information IM in association with text information Tx.
[0090] In the third embodiment, when the processor 21 of the decoding device 14 executes the decoding process to generate a decoded image Pd of the original image Po, the processor 21 inputs the text information Tx and auxiliary image acquired for the characteristic region Af and the text information Tx and auxiliary image acquired for the non-characteristic region An into the image generation model Mg to generate the decoded image Pd. This process can further improve the reproduction accuracy of the original image Po in the generated decoded image Pd.
[0091] Four modifications of the third embodiment have been described above, but embodiments that combine two or more of these modifications may also be envisioned.
[0092] <<Fourth Embodiment of the Present Invention>> In the above-described embodiment, a decoded image Pd of an original image Po is generated using an image generation model Mg (i.e., image generation AI) constructed by learning, but this model may be updated as appropriate. In other words, re-learning for constructing the image generation model may be performed to update the image generation model Mg. This case is defined as the fourth embodiment of the present invention, and the fourth embodiment will be described below. Note that in the fourth embodiment, differences from the first to third embodiments will be mainly described, and descriptions of commonalities with these embodiments will be omitted.
[0093] In the fourth embodiment, similar to the first to third embodiments, text information Tx is input to an image generation model Mg in a decoding process to generate a decoded image Pd of an original image Po. Here, the decoded image Pd is not necessarily identical to the original image Po. Therefore, in the fourth embodiment, one or more processors 21 included in the image processing system 10 further execute a first update process after executing the decoding process, in which re-learning is performed based on the original image Po and the decoded image Pd to update the image generation model Mg. This updates the image generation model Mg so that the reproduction accuracy of the original image Po for the decoded image Pd is improved; more specifically, it is possible to reset each parameter (hyperparameter) within the model. Note that the information used in the re-learning to update the image generation model Mg, i.e., the learning data, may include the original image Po and the decoded image Pd, as well as the text information Tx used to generate the decoded image Pd.
[0094] Furthermore, in the fourth embodiment, for the reasons for realizing the above-described configuration, the text conversion process, the decoding process, and the first update process may be executed by a single computer. That is, the image processing system 10 according to the fourth embodiment may be configured by a single computer. Specifically, an application program for implementing the functions of the text conversion device 12 and an application program for implementing the functions of the decoding device 14 may be installed on the same PC. However, even in the fourth embodiment, the image processing system 10 may be configured by multiple computers, and the text conversion device 12 and the decoding device 14 may each be configured by separate computers.
[0095] 12, the image processing flow according to the fourth embodiment is the same as that according to the first embodiment, except for the addition of step S025, which executes a first update process (S021 to S024, S026). In the first update process, as described above, the processor 21 of the computer constituting the image processing system 10 performs re-learning based on the original image Po and the decoded image Pd to update the image generation model Mg (S025). After executing the first update process, the processor 21 also executes a saving process (S026), and associates information about the updated image generation model Mg (model information IM) with text information Tx and saves them in a predetermined storage location.
[0096] [Variation of the Fourth Embodiment] In the above-described fourth embodiment, the image generation model Mg (image generation AI) is updated by performing re-learning based on the original image Po and the decoded image Pd. However, the target of the update may also be the text generation model Mt. That is, as a variation of the fourth embodiment, an embodiment in which re-learning for constructing a text generation model is performed to update the text generation model Mt may be considered. In such a variation, one or more processors 21 included in the image processing system 10 further perform a second update process after performing the decoding process. In the second update process, the processor 21 performs re-learning based on the original image Po, the decoded image Pd, and the text information Tx used to generate the decoded image Pd, thereby updating the text generation model Mt. This allows the text generation model Mt to be updated so that the reproduction accuracy of the original image Po is improved for the decoded image Pd generated from the text information Tx; more specifically, each parameter (hyperparameter) within the model can be reset.
[0097] The image processing flow according to the above-described modified example is the same as that according to the first embodiment (S031 to S034, S036), except for the addition of step S035 for executing a second update process, as shown in Fig. 13. In the second update process, as described above, the processor 21 of the computer constituting the image processing system 10 updates the text generation model Mt by re-learning based on the text information Tx, the original image Po, and the decoded image Pd (S035).
[0098] In the fourth embodiment and its modifications, both the image generation model Mg and the text generation model Mt may be retrained to update these two learning models. That is, the processor 21 of the computer constituting the image processing system 10 may execute both the first update process and the second update process.
[0099] <<Fifth Embodiment of the Present Invention>> As a further modification of the present invention, a case can be considered in which acquisition-related information of the original image Po is used in the text conversion process and the decoding process. Such a case is referred to as the fifth embodiment of the present invention, and the fifth embodiment will be described below. Note that in the fifth embodiment, differences from the first to fourth embodiments will be mainly described, and a description of points in common with these embodiments will be omitted.
[0100] In the fifth embodiment, when the processor 21 of the text conversion device 12 receives an input of an original image Po, it executes an information identification process to identify acquisition-related information for the original image Po. The acquisition-related information is information related to the acquisition of the original image Po and is information for identifying the conditions under which the original image Po was acquired. More specifically, if the original image Po is a photographed image, the acquisition-related information for the original image Po includes information on the date and time when the photographed image was acquired (photographed), location information of the photographed location, information on the climate and weather at the time of photographing, and information obtainable through the functions of a photographing device such as a camera. Information obtainable through the functions of a photographing device such as a camera may include, for example, the illuminance at the time of photographing, the focus position, and the device specifications. Furthermore, the acquisition-related information for the original image Po may include audio information obtained by collecting audio generated at the time of photographing the original image Po or during a time period including the time of photographing using a microphone mounted on the photographing device. The photographed image serving as the original image Po may be a still image or a video. Furthermore, the acquired related information may include not only information at the time of shooting of the photographed image, but also information at at least one of the time points immediately before and immediately after shooting.
[0101] On the other hand, if the original image is a drawn image such as a CG image, information related to the acquisition of the original image Po includes information on the date and time the image was drawn, information on the attributes of the artist (specifically, age, gender, etc.), and information on the drawing tools such as software used to draw the image.
[0102] The procedure for identifying the acquisition-related information of the original image Po is not particularly limited, but for example, the acquisition-related information may be identified by referring to or analyzing incidental information (tag information) of the original image Po and information associated with the original image Po. In the following description of the fifth embodiment, it is assumed that the original image Po is a photographed image.
[0103] In the fifth embodiment, after the acquisition-related information of the original image Po is identified, the processor 21 of the text conversion device 12 executes a text conversion process. In this process, the processor 21 outputs text information Tx based on the original image Po and the acquisition-related information IP of the original image Po (strictly speaking, information related to the capture of the original image Po).
[0104] Specifically, as shown in FIG. 14 , the processor 21 inputs the original image Po and the acquired related information IP into the text generation model Mt and outputs text information Tx about the original image Po. The text information Tx output in this case is more detailed than text information Tx output based only on the original image Po. For example, if the acquired related information IP indicates that the original image Po was captured during the daytime, the text information Tx is output under the assumption that the original image Po is likely to have been captured during the daytime. Furthermore, if the acquired related information IP is sound information about waves, keywords related to "waves," such as a beach, may be included in the text information Tx, or the priority of keywords related to "waves" may be increased. Keywords with higher priority are more likely to be included in the text information Tx.
[0105] In the fifth embodiment, a decoding process is performed using the text information Tx output based on the original image Po and its acquisition-related information IP, to generate a decoded image Pd of the original image Po. The generated decoded image Pd has a higher reproduction accuracy of the original image Po than a decoded image Pd generated using the text information Tx output based only on the original image Po.
[0106] 14, the decoding process of the fifth embodiment may use text information Tx and acquisition-related information IP of the original image Po. That is, in the decoding process, for example, acquisition-related information IP relating to the time when the original image Po was captured or the sound collected at the time of capture may be input to the image generation model Mg along with the text information Tx to generate a decoded image Pd of the original image Po. The generated decoded image Pd will have a higher reproduction accuracy of the original image Po than a decoded image Pd generated based only on the text information Tx.
[0107] 14 , in the text conversion process, text information Tx is output based on the original image Po and its acquisition-related information IP, and in the decoding process, a decoded image Pd is generated based on the text information Tx and the acquisition-related information IP. However, the present invention is not limited to this, and the acquisition-related information IP may be used in either the text conversion process or the decoding process.
[0108] <<Sixth Embodiment of the Present Invention>> In the above-described embodiments, the target of the text conversion process is a single image (particularly a still image), but cases in which the text conversion process is performed on an image group consisting of multiple images can also be considered. Such a case is referred to as the sixth embodiment of the present invention, and the sixth embodiment will be described below. Note that in the sixth embodiment, differences from the first to fifth embodiments will be mainly described, and descriptions of points in common with these embodiments will be omitted.
[0109] In the sixth embodiment, a group of images is input as original images Po, and the processor 21 of the text conversion device 12 accepts the input of the group of images in an input acceptance process. Here, the group of images is a moving image consisting of a plurality of frame images. In this case, the processor 21 performs text conversion processing on the plurality of frame images included in the input moving image using either the first pattern or the second pattern.
[0110] The pattern to be used for the text conversion process may be determined based on a selection operation by the user. That is, the processor 21 of the text conversion device 12 may accept a selection operation by the user and perform the text conversion process based on the pattern specified by the selection operation.
[0111] When performing text conversion processing using the first pattern, text information Tx representing each frame image in the video is output for each frame image, as shown in Fig. 15. In this case, text information Tx is obtained for each frame image in the video, so that when each frame image is decoded, the text information Tx corresponding to that frame image can be used to generate an appropriate decoded image for each frame image.
[0112] On the other hand, when the text conversion process is performed using the second pattern, as shown in FIG. 16 , text information Tx is output for a first image in a video using the normal text conversion process procedure. For a second image in a video, text information Tx is output, describing differences between the second image and the first image. Here, the first image is a reference frame image among multiple frame images included in the video, and the second image is a frame image acquired (captured) later in the video than the first image. For example, as shown in FIG. 16 , the first frame image in the video, i.e., the frame image captured immediately after the start of video capture, may be designated as the first image, and subsequent frame images may be designated as the second images. Furthermore, of two consecutive frame images, the frame image captured earlier may be designated as the first image, and the frame image captured later may be designated as the second image. Furthermore, in a case where the first frame image is designated as the first image and subsequent frame images are designated as second images, if a predetermined condition is met during the video, the immediately following frame image may be designated as the new first image, and subsequent frame images may be designated as second images. The predetermined condition may be a significant change in the subject in the frame image, such as a change in the subject's facial expression, position in the image, size relative to the image, etc. The predetermined condition may also be the passage of a certain amount of time, in which case the first image will be switched each time the certain amount of time has passed.
[0113] In the text generation process executed in the second pattern, when outputting text information Tx expressing the differences between the second image and the first image, it is advisable to compare the first image and the second image using a known image analysis method and identify areas in the second image that have changed from the first image, i.e., differences (differences).The identified differences can then be input into a text generation model Mt, whereby text information Tx expressing the differences can be output.
[0114] As described above, when the text conversion process is performed using the second pattern, for a frame image corresponding to the second image, text information Tx that represents the differences from the first image is output, rather than text information Tx that represents the entire image. This makes it possible to further reduce the volume of the text information Tx for the second image, thereby further improving the compression efficiency when the second image is converted to text (encoded) and compressed.
[0115] [Variation of the Sixth Embodiment] In the sixth embodiment described above, a video consisting of a plurality of frame images captured consecutively at a fixed frame interval is used as the original image Po to perform the text conversion process. However, the configuration of the sixth embodiment can also be applied to cases where a group of images other than a video is used as the original image Po to perform the text conversion process. An example of a group of images other than a video is a group of images obtained by capturing multiple images at different times related to a certain shooting theme, specifically, a group of images obtained by capturing images of a certain event multiple times with different subjects, angles of view, shooting methods, etc.
[0116] <<Seventh Embodiment of the Present Invention>> In the second embodiment, the third embodiment, and their modifications (particularly, the third embodiment to the third embodiment), the generation process is performed on a single image (particularly, a still image). However, cases in which the generation process is performed on an image group consisting of multiple images can also be considered. Such a case is referred to as the seventh embodiment of the present invention, and the seventh embodiment will be described below. Note that the seventh embodiment will mainly describe the differences from the second embodiment, the third embodiment, and their modifications, and a description of the points in common with these embodiments will be omitted.
[0117] In the seventh embodiment, a group of images is input as original images Po, and the processor 21 of the text conversion device 12 accepts the input of the group of images in the input acceptance process. Here, the group of images is a moving image consisting of a plurality of frame images. In this case, it is not necessary to save reduced images Ps or extracted images Pe of characteristic regions Af for all frame images included in the moving image. Therefore, in the seventh embodiment, the processor 21 switches the execution pattern of the text conversion process and generation process between the third pattern and the fourth pattern depending on the frame images included in the moving image.
[0118] When the execution pattern is the third pattern, the processor 21 of the text conversion device 12 executes both the text conversion process and the generation process on the frame image to be processed. That is, as shown in Fig. 17, for the frame image to be processed when the execution pattern is the third pattern (for example, frame images #1 and #i in Fig. 17), text information Tx representing the frame image is output, and a reduced image Ps of the frame image is generated.
[0119] When the execution pattern is the fourth pattern, the processor 21 of the text conversion device 12 performs the text conversion process on the frame image to be processed, but does not perform the generation process. In other words, as shown in Figure 17, for the frame image to be processed when the execution pattern is the fourth pattern (for example, frame image #2 in Figure 17), text information Tx representing the frame image is output, but a reduced image Ps of the frame image is not generated.
[0120] It should be noted that which of the third and fourth execution patterns to apply to each frame image can be determined based on its relationship to a reference frame image in the video. The reference frame image is, for example, the frame image immediately after video capture begins, i.e., frame image #1, and the third pattern is applied to the reference frame image. On the other hand, among frame images #2 and onward, the third pattern is applied to frame images that satisfy a predetermined condition, and the fourth pattern is applied to other frame images. The predetermined condition may be that the subject in the frame image changes significantly from the reference frame image, for example, that the subject's facial expression, position in the image, size relative to the image, etc. change. Furthermore, the predetermined condition may be that a feature amount of each frame image is identified, and the difference between the identified feature amount and the feature amount of the reference frame image is equal to or greater than a threshold.
[0121] As described above, in the seventh embodiment, whether or not to execute the generation process is set for each of the multiple frame images included in a moving image, and reduced images Ps of only those frame images for which the generation process is required are saved. This reduces the processing load compared to saving reduced images Ps for all frame images included in a moving image.
[0122] As an example of the seventh embodiment, a case has been given in which, of multiple frame images included in a moving image, a reduced image Ps of that frame image is generated for which the third pattern is applied, but the present invention is not limited to this. For a frame image for which the third pattern is applied and the generation process is executed, a characteristic region Af in that frame image may be identified, and an extracted image Pe of the characteristic region Af may be generated, or related information related to the characteristic region Af may be acquired.
[0123] [Modification of Seventh Embodiment] In the seventh embodiment described above, a moving image made up of a plurality of frame images captured consecutively at a fixed frame interval is used as the original image Po, and the execution pattern of the text conversion process and the generation process is switched depending on the frame images in the moving image. However, the configuration of the seventh embodiment can also be applied to cases where a group of images other than a moving image is used as the original image Po. As with the modification of the sixth embodiment, an example of a group of images other than a moving image is a group of images obtained by capturing multiple images at different times regarding a certain shooting theme.
[0124] <<Other Embodiments>> Specific embodiments of the present invention have been described above, but the above-described embodiments are merely examples given to facilitate understanding of the present invention and are not intended to limit the present invention. That is, the present invention may be modified or improved from the embodiments described below without departing from the spirit of the present invention. Furthermore, the present invention includes equivalents thereof. Furthermore, embodiments of the present invention may include a combination of the above-described embodiments with one or more of the following modifications.
[0125] (Regarding the Text Generation Means and the Image Generation Means) In the above-described embodiment, a text generation model Mt (text generation AI) constructed through learning is used as the text generation means, but this is not limited to this. For example, correspondences between image features and corresponding words and phrases may be specified in advance and stored as table data. Then, the original image Po may be analyzed to specify features of each region in the original image Po, words and phrases corresponding to the specified features may be searched for in the table data, and the words and phrases searched for in each region of the original image Po may be compiled to generate text information Tx based on the original image Po.
[0126] Furthermore, in the above-described embodiment, an image generation model Mg (image generation AI) constructed through learning is used as the image generation means, but this is not limited thereto. For example, correspondences between words and corresponding images (more specifically, images of people or objects) may be specified in advance and stored as table data. Then, the text information Tx may be analyzed to identify words in the text information Tx, images corresponding to the identified words and words may be searched for in the table data, and the images searched for for each word in the text information Tx may be integrated to generate a decoded image Pd based on the text information Tx. Furthermore, the table data may be stored as information necessary for executing the decoding process, associated with the text information Tx used in the decoding process.
[0127] (Regarding Correction of Text Information) In the above embodiment, the text conversion device 12 converts (encodes) the original image Po into text using the text generation model Mt and outputs the text information Tx. However, the output text information Tx may be corrected at the user's discretion. In other words, one or more processors 21 of the image processing system 10, specifically the processor of the text conversion device 12, may further execute a correction process to correct the output text information Tx (strictly speaking, the text content indicated by the text information Tx) based on a user operation. In this case, after the user checks the text information Tx output by the text generation model Mt, the user can supplement the text information Tx with missing information or change incorrect information.
[0128] (Regarding the Computers Constituting the Image Processing System) In the above embodiment, the image processing system 10 is composed of multiple computers, and the computer constituting the text converter 12 and the computer constituting the decoding device 14 exist separately. However, this is not limited to this, and the text converter 12 and the decoding device 14 may be composed of the same computer. For example, a single computer may be equipped with the functions of both the text converter 12 and the decoding device 14. In other words, the image processing system 10 may be composed of a single computer. As a specific example, a server computer for a cloud service may function as both the text converter 12 and the decoding device 14. In this case, a user transmits an original image Po from their own PC (client terminal) to the server computer, and the server computer accepts the input of the original image Po and outputs text information Tx about the original image Po. The user can receive the output text information Tx from the server computer via the network N. Furthermore, to decode the original image Po from the text information Tx, the user transmits the text information Tx from their own PC to the server computer. As a result, the server computer receives the text information Tx and generates a decoded image Pd of the original image Po based on the text information Tx. The generated decoded image Pd is sent from the server computer to the user's PC via the network N. This allows the user to use the decoded image Pd of the original image Po.
[0129] The image processing system 10 may also be configured with a photographing device such as a digital camera having image processing and communication functions. In this case, each time an image is captured, the photographing device inquires of the user whether or not the captured image needs to be converted to text (compressed). If the user replies that the image needs to be converted to text (compressed), the photographing device (more specifically, the processor 21 of the photographing device) may perform a text conversion process on the captured image as the original image Po and output text information Tx. The photographing device may also attach information for decoding to the text information Tx; specifically, it may store information about the image generation model Mg (model information IM) in association with the text information Tx.
[0130] (Regarding the Processor Configuration) The image processing system of the present invention includes at least one or more processors, which may include various types of processors. The various processors include, for example, a CPU, which is a general-purpose processor that executes software (programs) and functions as various processing units. Furthermore, each process included in the image processing of the present invention may be executed by any computer. Furthermore, the any computer may execute these processes using a processor, a program, or a combination thereof. The any computer may be a system such as a general-purpose computer, a computer for a specific application, a workstation, or other hardware element capable of executing a program.
[0131] The processor may be configured with one or more pieces of hardware, and the type of hardware is not limited. For example, the processor may be configured with hardware such as a central processing unit (CPU), a micro processing unit (MPU), a programmable logic device such as a field programmable gate array (FPGA), a dedicated circuit for executing specific processing such as an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or a neural processing unit (NPU).
[0132] Furthermore, the processor has each unit or each means that executes various processes in the present invention. Furthermore, the type of hardware may be a combination of different types of hardware. When multiple pieces of hardware are configured to execute one or more processes of a certain processor, the multiple pieces of hardware may exist in devices that are physically separate from each other, or may exist in the same device. Furthermore, in any embodiment of the present invention, the order of each process performed by the processor is not limited to the order described above and may be changed as appropriate. The hardware is configured by an electric circuit (circuitry) that combines circuit elements such as semiconductor elements.
[0133] Furthermore, each embodiment of the present invention may be realized by hardware, software, firmware, microcode, or a combination thereof. Software, firmware, and microcode are configured by a program. A program may also be, for example, a group of program modules, each function of which may be implemented by a processor configured to perform the respective function. The program may also be program code and multiple code segments stored in one or more non-transitory computer-readable media (e.g., storage media or other storages). The program may also be stored in multiple non-transitory computer-readable media that reside in physically separate devices. Program code or code segments may represent procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Program code or code segments may be connected to other code segments or hardware circuits by sending or receiving information, data, arguments, parameters, or memory contents.
[0134] 10 Image processing system 12 Text conversion device 14 Decoding device 21 Processor 22 Memory 23 Communication interface 24 Storage 25 Input device 26 Output device Af Feature region An Non-feature region IM Model information IP Acquisition related information Mg Image generation model Mgi Input layer Mgm Hidden layer Mgo Output layer Mt Text generation model Mti Input layer Mtm Hidden layer Mto Output layer N Network Pe Extracted image Po Original image Pd Decoded image Ps Reduced image Tx Text information
Claims
1. An image processing system having a processor, wherein the processor executes an input reception process for receiving input of an original image, a text conversion process for outputting text information relating to text expressing the original image based on the original image, an information reception process for receiving the text information, and a decoding process for generating a decoded image of the original image based on the text information.
2. A text conversion device comprising the image processing system according to claim 1, the processor being included in the processor, and having a processor that executes the input acceptance process and the text conversion process.
3. A decoding device that constitutes the image processing system according to claim 1, that is included in the processor, and that executes the information reception process and the decoding process.
4. The image processing system according to claim 1, wherein the processor further executes a generation process for generating a reduced image of the original image, an extracted image of a partial area of the original image, or related information relating to a partial area of the original image.
5. An image processing system as described in claim 4, wherein in the decoding process, the processor generates the decoded image based on the text information and the reduced image, the extracted image or the related information generated in the generation process.
6. The image processing system according to claim 1, wherein the processor further executes a storage process for storing information required for executing the decoding process in association with the text information.
7. The image processing system of claim 6, wherein in the decoding process, the processor inputs the text information into an image generation model constructed by learning to generate the decoded image, and the information required to execute the decoding process is information of the image generation model.
8. An image processing system according to claim 4, wherein the processor further executes a storage process for storing the reduced image, the extracted image or the related information generated in the generation process in association with the text information.
9. The image processing system of claim 1, wherein in the decoding process, the processor inputs the text information into an image generation model constructed by learning to generate the decoded image, and the processor further executes a first update process in which re-learning is performed based on the original image and the decoded image to update the image generation model.
10. The image processing system of claim 1, wherein in the text conversion process, the processor inputs the original image into a text generation model constructed by learning and outputs the text information, and the processor further executes a second update process in which re-learning is performed based on the text information, the original image, and the decoded image to update the text generation model.
11. The image processing system of claim 1, wherein in the text conversion process, the processor outputs the text information based on the original image and information related to the acquisition of the original image.
12. The image processing system of claim 11, wherein the original image is a photographed image, and in the text conversion process, the processor outputs the text information based on the original image and information related to the photographing of the original image.
13. The image processing system of claim 1, wherein, when the original image is a photographed image, in the decoding process, the processor generates the decoded image based on the text information and information related to the photographing of the original image.
14. The image processing system of claim 1, wherein, when the original image is a photographed image, in the text conversion process, the processor outputs the text information based on the original image and information related to the photographing of the original image, and in the decoding process, the processor generates the decoded image based on the text information and information related to the photographing of the original image.
15. The image processing system of claim 1, wherein, when the processor receives input of an image group consisting of a plurality of images in the input reception process, the processor performs the text conversion process in a first pattern in which the text information representing an image included in the image group is output for each image, or in a second pattern in which, when outputting the text information representing a second image acquired later in time than a first image included in the image group, the text information representing the differences between the second image and the first image is output.
16. The image processing system of claim 4, wherein when the processor receives input of an image group including a plurality of images in the input reception process, the processor switches the execution pattern of the text conversion process and the generation process between a third pattern in which the text conversion process and the generation process are executed, and a fourth pattern in which the text conversion process is executed but the generation process is not executed, depending on the images included in the image group.
17. An image processing system according to claim 15 or 16, wherein the image group is a video consisting of a plurality of frame images.
18. The image processing system according to claim 1, wherein the processor further executes a correction process for correcting the text information based on a user operation.
19. An image processing method in which a processor performs the following steps: accepting input of an original image; outputting text information regarding text representing the original image based on the original image; accepting the text information; and generating a decoded image of the original image based on the text information.
20. The image processing method according to claim 19, further comprising the step of generating a reduced image of the original image, an extracted image of a partial region of the original image, or related information relating to a partial region of the original image, by the processor.
21. A program for causing a processor to execute each step included in the image processing method according to claim 19 or 20.
22. A processor-readable recording medium having recorded thereon a program for causing a processor to execute each step included in the image processing method according to claim 19 or 20.
Citation Information
Patent Citations
Method and apparatus for training image caption model, and storage medium
US20210034981A1
Devices and methods for providing images and image capturing based on a text and providing images as a text
WO2021170230A1
Cited By
Image transmission and reception system, transmitter, receiver, computer program, and image transmission and reception method
JP7838169B1