Image processing device, image processing method, and program

The image processing device addresses the challenge of high data storage by separating the main subject and background, generating a prompt for the background, and using trained models to reconstruct the image, achieving reduced data storage with minimal quality loss.

JP2025187999APending Publication Date: 2025-12-25CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025074897
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-14
Filing Date
2025-04-28
Publication Date
2025-12-25

AI Technical Summary

Technical Problem

Existing image processing techniques fail to reduce the amount of data required to store images, particularly when both the image and output tensor are output, leading to increased storage requirements.

Method used

An image processing device that separates the main subject from the background in an input image, generates text data indicating background characteristics as a prompt, and uses trained models to restore the background, thereby reducing the data amount by converting the background image into a prompt and generating a restored image with reduced data.

Benefits of technology

The method effectively reduces the data required to store images while maintaining image quality by separating the main subject and background, generating a prompt for the background, and using trained models to reconstruct the image with minimal loss of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025187999000001_ABST
    Figure 2025187999000001_ABST
Patent Text Reader

Abstract

To solve the problem in which the total data volume required to store images increases.SOLUTION: An image processing device includes determination means for determining a main subject in an input image containing the main subject and a background, and generating from the input image a first image containing the main subject and a second image containing the background, and generation means for generating text data indicating features of the background as a prompt on the basis of the second image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an image processing device, an image processing method, and a program. [Background technology]

[0002] Patent Document 1 discloses a technology for inputting an image into a trained model such as a deep neural network (DNN) to perform recognition processing and generate metadata such as an output tensor. This technology outputs at least one of the image and the output tensor. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-18997 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the above-mentioned technique cannot reduce the amount of image data by the output tensor. In particular, when the above-mentioned technique outputs the image and the output tensor, the total amount of data required to store the image increases.

[0005] Therefore, the present disclosure provides a technique that can reduce the amount of data required to store images. [Means for solving the problem]

[0006] In order to solve this problem, for example, an image processing device according to the present disclosure has the following configuration: a determining means for determining a main subject of an input image including a main subject and a background, and generating a first image including the main subject and a second image including the background from the input image; generating means for generating text data indicating characteristics of the background as a prompt based on the second image; Equipped with. [Effects of the Invention]

[0007] According to the present disclosure, the amount of data required to store images can be reduced. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 2 is a block diagram showing the hardware configuration and functional configuration of a control system of the image processing apparatus according to the first embodiment. [Figure 2] FIG. 3 is a flowchart showing image processing according to the first embodiment. [Figure 3] 5A to 5C are diagrams illustrating the image processing process and prompting in the image processing according to the first embodiment. [Figure 4] FIG. 4 is a flowchart showing a restoration process according to the first embodiment. [Figure 5] 5A to 5C are diagrams illustrating the image processing steps of the restoration processing according to the first embodiment. [Figure 6] FIG. 10 is a block diagram showing the functional configuration of an image processing apparatus according to a second embodiment. [Figure 7] FIG. 10 is a flowchart showing image processing according to the second embodiment. [Figure 8] 10A to 10C are diagrams illustrating the image processing process and prompting in the image processing according to the second embodiment. [Figure 9] FIG. 10 is a flowchart showing a restoration process according to the second embodiment. [Figure 10] 10A to 10C are diagrams illustrating the image processing steps of the restoration processing according to the second embodiment. [Figure 11] FIG. 10 is a block diagram showing the functional configuration of an image processing apparatus according to a third embodiment. [Figure 12] FIG. 11 is a flowchart showing image processing according to the third embodiment. [Figure 13] 10A to 10C are diagrams illustrating the image processing steps and prompts of the image processing according to the third embodiment. [Figure 14] FIG. 11 is a flowchart showing a restoration process according to the third embodiment. [Figure 15] 10A to 10C are diagrams illustrating the image processing steps of the restoration processing according to the third embodiment. [Figure 16] FIG. 10 is a block diagram showing the functional configuration of an image processing apparatus according to a fourth embodiment. [Figure 17] FIG. 10 is a flowchart showing image processing according to the fourth embodiment. [Figure 18] 13A to 13C are diagrams illustrating the moving image processing process and prompting in the image processing according to the fourth embodiment. [Figure 19] FIG. 13 is a flowchart showing a restoration process according to the fourth embodiment. [Figure 20] 13A to 13C are diagrams illustrating the process of moving image restoration processing according to the fourth embodiment. [Figure 21] FIG. 11 is a block diagram showing the functional configuration of an image processing apparatus according to a fifth embodiment. [Figure 22] FIG. 13 is a flowchart showing image processing according to the fifth embodiment. [Figure 23] FIG. 13 is a flowchart showing a restoration process according to the fifth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0010] First Embodiment An image processing device in a first embodiment will be described with reference to the drawings. Fig. 1 is a block diagram showing the hardware configuration and functional configuration of a control system of the image processing device in the first embodiment. Fig. 1(a) is a block diagram showing the hardware configuration of the image processing device.

[0011] The image processing device 100 may be an information processing device (also called a computer) such as a personal computer, a smartphone, or a tablet terminal, and is not particularly limited.

[0012] 1(a), the image processing device 100 includes a processor 101, a nonvolatile memory 102, a memory 103, an input IF 113, an output IF 111, a communication IF 112, and a system bus 105. The processor 101, the nonvolatile memory 102, the memory 103, the input IF 113, the output IF 111, and the communication IF 112 are connected by the system bus 105 so as to be able to transmit and receive data to and from each other.

[0013] The processor 101 is an arithmetic processing device that processes various calculations. The processor 101 may be, for example, a CPU (Central Processing Unit). Instead of or in addition to the CPU, the processor 101 may include other processors such as an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), or a QPU (Quantum Processing Unit). The processor 101 realizes various functions by executing programs stored in the nonvolatile memory 102. In this way, the processor 101 controls the operation of each unit (each functional block) of the image processing device 100 via the system bus 105. The processor 101 may use the memory 103 as a work memory in accordance with the programs stored in the nonvolatile memory 102, for example.

[0014] The nonvolatile memory 102 is a nonvolatile memory that can be electrically erased and stored. The nonvolatile memory 102 may be, for example, a read-only memory (ROM), a hard disk drive (HDD), or a solid state drive (SSD). The nonvolatile memory 102 stores data such as computer programs (also referred to as programs), parameters required for executing the programs, and image data to be processed by the programs. Note that the term "image" may include both images and image data. For example, the nonvolatile memory 102 stores an input image, various programs for operating the processor 101, a first trained model, a second trained model, a third trained model, and image processing results.

[0015] In this embodiment, the input image is stored in advance in the non-volatile memory 102, but this is merely an example. For example, the input image may be input from an external storage device or a network via the communication IF 112 or the like. There are no particular limitations on the method for inputting the input image. When the input image is input from outside, it is not necessary to store the input image.

[0016] The memory 103 is a volatile memory that can read and write data at high speed, and may be a RAM such as a dynamic random access memory (DRAM).

[0017] The input IF 113 is an interface that receives input from an external input device. For example, the input IF 113 receives user input from an input device such as a mouse or keyboard, and outputs the input to the processor 101.

[0018] The output IF 111 outputs data to an external device. For example, the output IF 111 outputs image data to an image display device such as a display to display an image based on an instruction from the processor 101. The output IF 111 may also output data such as image data to an external storage device such as an HDD or SSD to store the data.

[0019] The communication IF 112 is an interface connected to a network such as a wide area network (WAN) or a local area network (LAN). The communication method of the communication IF 112 may be either wireless communication or wired communication, and is not particularly limited. The communication IF 112 realizes sending and receiving of data to and from an external device via the network.

[0020] Fig. 1(b) is a functional block diagram illustrating the functions of the image processing device. As shown in Fig. 1(b), the image processing device 100 has a main subject determination unit 104, a prompt generation unit 106, an input unit 107, a background generation unit 108, and an image restoration unit 109. The main subject determination unit 104 is an example of a determination means. The prompt generation unit 106 is an example of a generation means. The background generation unit 108 and the image restoration unit 109 are examples of a restoration generation means.

[0021] Some or all of the functions of main subject determination unit 104, prompt generation unit 106, input unit 107, background generation unit 108, and image restoration unit 109 are realized by processor 101 reading out a program stored in nonvolatile memory 102, expanding it in memory 103, and executing it. Furthermore, some or all of the functions of main subject determination unit 104, prompt generation unit 106, input unit 107, background generation unit 108, and image restoration unit 109 may be realized by one or more circuits, such as an ASIC (Application Specific Integrated Circuit) and a PLD (Programmable Logic Device) including an FPGA (Field Programmable Gate Array).

[0022] The main subject determination unit 104 determines the main subject in the input image through calculation processing using a first trained model. Based on the determination result, the main subject determination unit 104 separates the input image into a main subject image including the main subject and a background image including the background, and generates a main subject image and a background image from the input image. The main subject image is an example of a first image, and the background image is an example of a second image. The main subject determination unit 104 stores the separated main subject image, image size information of the input image, and position information of the main subject in the input image in non-volatile memory 102. In this embodiment, the main subject determination unit 104 stores each piece of information in non-volatile memory 102. However, the main subject determination unit 104 may output and store the information in an external storage device or a storage device on a network via the output IF 111 or the communication IF 112. In this embodiment, the image size information of the input image, the position information of the main subject in the input image, and the storage destination of the main subject image are not particularly limited. The main subject determination unit 104 also temporarily stores the background image in memory 103. The main subject determination unit 104 may determine the main subject using other image processing techniques such as pattern matching in addition to neural network calculation processing.

[0023] The prompt generation unit 106 generates text data indicating background features in a background image as a prompt through arithmetic processing using the second trained model. The prompt generation unit 106 saves the generated prompt indicating the background features in the non-volatile memory 102. The prompt generation unit 106 saves the prompt in the non-volatile memory 102, but may also output and save the prompt to an external storage device or a network via the output IF 111 or the communication IF 112. This embodiment does not particularly limit the destination where the generated prompt indicating the background features is saved. The prompt generation unit 106 may generate the prompt indicating the background features using other image processing techniques, such as pattern matching, in addition to neural network arithmetic processing.

[0024] The input unit 107 accepts input of instructions from a user. The input unit 107 accepts input via the input IF 113 from, for example, a mouse, a keyboard, a touch panel, or the like. The input unit 107 stores, for example, the amount of output data input by the user as an instruction in the memory 103. The input unit 107 may have a display device for prompting the user to input instructions. The display device may be integrated with, for example, a touch panel.

[0025] The background generation unit 108 inputs a prompt, which is text data indicating the characteristics of the background, into a third trained model and performs arithmetic processing to generate a generated background image by restoring the background. The background generation unit 108 temporarily stores the generated background image in memory 103. The third trained model may be a single trained model or a combination of multiple trained models, such as a VLM (Visual Language Model) and a diffusion model. An example of the third trained model may be DALL·E2 developed by OpenAI (registered trademark).

[0026] Image restoration unit 109 generates a restored image by restoring the input image using the generated background image and the main subject image. Specifically, image restoration unit 109 generates the restored image by superimposing the main subject image at a position on the generated background image that corresponds to the position information of the main subject in the input image. Image restoration unit 109 stores the generated restored image in non-volatile memory 102. Image restoration unit 109 may output the generated restored image to an external display device or the like via output IF 111 for display.

[0027] Although the nonvolatile memory 102 has been described as storing the first to third trained models, this is not limiting. For example, if neural network arithmetic processing is not used to determine the main subject in the main subject determination unit 104, the nonvolatile memory 102 does not need to store the first trained model. If neural network arithmetic processing is not used to generate a prompt indicating background characteristics in the prompt generation unit 106, the nonvolatile memory 102 does not need to store the second trained model.

[0028] Next, a description will be given of image processing from separating the main subject and the background from the input image to making the background prompt in the first embodiment. Fig. 2 is a flowchart of the image processing in the first embodiment.

[0029] Before image processing begins, an input image to be used by the main subject determination unit 104, a first trained model for determining the main subject, and a second trained model for generating text data indicating the characteristics of the background image to be used by the prompt generation unit 106 are pre-stored in the non-volatile memory 102.

[0030] In S200, the input unit 107 acquires from the user an output data amount as a threshold value for the total data amount of the main subject image data and the text data amount indicating the background characteristics. The input unit 107 stores the acquired output data amount in the memory 103.

[0031] Next, in S201, main subject determination unit 104 determines the main subject of the input image, separates the main subject from the background, and generates a main subject image and a background image from the input image. Specifically, main subject determination unit 104 reads the input image and the first trained model from non-volatile memory 102. Using the first trained model for the input image, main subject determination unit 104 determines the main subject in the input image through neural network calculation processing. Main subject determination unit 104 separates the input image into a main subject region image and a background image by cutting out the region in which the main subject is located in the input image with a shape (here, a rectangle) that fits the main subject. Thereafter, main subject determination unit 104 performs transparency processing on the portion other than the main subject in the main subject region image cut out in the rectangle, i.e., the background included in the main subject region image, to generate the main subject image. In other words, main subject determination unit 104 removes the background from the main subject region image to generate the main subject image. Main subject determination unit 104 stores the generated main subject image, image size information of the input image, and position information of the main subject in the input image in non-volatile memory 102. Main subject determination unit 104 also temporarily stores the background image in memory 103. Separation of the main subject image and the background image will be described in detail later.

[0032] Next, in S202, main subject determination unit 104 determines whether the data amount of the main subject image is less than the output data amount. Specifically, main subject determination unit 104 reads the output data amount received from the user from memory 103. Main subject determination unit 104 compares the data amount of the generated main subject image with the output data amount. If main subject determination unit 104 determines that the data amount of the main subject image is not less than the output data amount, it returns to S200 and prompts the user to input the output data amount again. This causes input unit 107 to acquire a new output data amount. Note that if S200 and subsequent steps are repeated, main subject determination unit 104 may omit S201. On the other hand, if main subject determination unit 104 determines that the data amount of the main subject image is less than the output data amount, it proceeds to S203.

[0033] Next, in S203, the prompt generation unit 106 converts the background image into a prompt. Specifically, the prompt generation unit 106 reads the background image from the memory 103 and also reads the second trained model from the non-volatile memory 102. The prompt generation unit 106 uses the second trained model for the background image to generate a prompt in which the features of the background image are converted into text through neural network calculation processing. Therefore, the prompt can also be considered text data. The prompt generation unit 106 stores a prompt indicating the features of the generated background image in the non-volatile memory 102. Here, the second trained model may generate a prompt indicating the features of the background based on the level of detail. The user may set the level of detail via the input unit 107. The higher the level of detail, the greater the amount of data in the prompt indicating the features of the background. A detailed description of converting a background image into a prompt will be given later.

[0034] Next, in S204, prompt generation unit 106 determines whether the total data amount of the main subject image and the prompt is less than the output data amount. Specifically, prompt generation unit 106 reads the output data amount received from the user from memory 103. Prompt generation unit 106 calculates the sum of the data amount of the main subject image read from non-volatile memory 102 and the data amount of the generated prompt indicating the characteristics of the background. Prompt generation unit 106 compares the calculated total data amount with the output data amount. If prompt generation unit 106 determines that the calculated total data amount is not less than the output data amount, the process proceeds to S205.

[0035] In S205, the prompt generation unit 106 reduces the level of detail of the prompt and proceeds to S203. The prompt generation unit 106 may reduce the level of detail based on a preset setting value. Alternatively, the prompt generation unit 106 may notify the user to reduce the level of detail and reduce the level of detail based on the level of detail input by the user.

[0036] In S203, the prompt generation unit 106 generates a prompt indicating the characteristics of the background by converting the background image into a prompt based on the lowered level of detail. Here, the prompt generated by the prompt generation unit 106 has a low level of detail, so the amount of data for the prompt is small.

[0037] Thereafter, prompt generating unit 106 repeats S205 and S203 until the total data amount of the main subject image and the prompt becomes smaller than the output data amount.

[0038] In S204, if prompt generating unit 106 determines that the total data amount of the main subject image and the prompt is less than the output data amount, the flow ends.

[0039] As a result, the image processing device 100 reduces the total data amount of the main subject image and the prompt data amount indicating the background characteristics from the output data amount received from the user, thereby reducing the data amount of the input image while suppressing a reduction in background information.

[0040] 3 is a diagram illustrating the image processing process and prompting in the image processing of the first embodiment. Next, the separation of the main subject image and the background image, which is the process in S201 of the flowchart, and the prompting of the background image, which is the process in S203, will be described in detail with reference to FIG.

[0041] FIG. 3(a) shows input image 300 of this embodiment. The image shows a man standing in front of a mountain lake on a clear day. As a result of determining the main subject of input image 300, main subject determination unit 104 determines the man located in the center as main subject 301. Main subject determination unit 104 cuts out the region of input image 300 where main subject 301, the man, is located, as a rectangular region that exactly fits main subject 301, as shown in FIG. 3(b), to generate main subject region image 305. For example, main subject determination unit 104 may generate main subject region image 305 by cutting out the rectangular region that fits main subject 301 using horizontal and vertical lines that pass through the top, bottom, left, and right edges of main subject 301 or points a predetermined distance outside the edges.

[0042] Main subject determination unit 104 generates main subject region image 305 including main subject 301 shown in Fig. 3(b), and then performs transparency processing on the portion other than main subject 301, i.e., the slight remaining background portion. As a result, main subject determination unit 104 generates main subject image 306, which is an image of main subject 301 shown in Fig. 3(c). In conjunction with the generation of main subject region image 305 or main subject image 306, main subject determination unit 104 generates and saves position information of main subject image 306 in input image 300.

[0043] The main subject determination unit 104 generates the remaining image after cutting out the main subject region image 305 from the input image 300 as a background image 307, which is an image of the background 302 shown in FIG. 3(d).

[0044] As a result, the main subject determination unit 104 separates the input image shown in Figure 3(a) into a main subject image 306, which is an image of the main subject 301 shown in Figure 3(c), and a background image 307, which is an image of the background 302 shown in Figure 3(d).

[0045] Next, the prompt generating unit 106 generates text data indicating characteristics from the image of the background 302 shown in FIG. 3(d) as a prompt 303 shown in FIG. 3(e).

[0046] As shown in FIG. 3(e), the prompt generation unit 106 generates prompt 303, which is text data, based on background image 307 and includes characteristics indicating the background situation, characteristics of objects present in the background, and the relative positions of the objects. Specifically, based on background image 307 shown in FIG. 3(d), the prompt generation unit 106 generates prompt 303 including overall background characteristics such as daytime and sunny weather, and partial characteristics such as three gently sloping mountains, a forest, a lake, and low grass, along with their relative positions. Because prompt 303 is text data, its data volume is very small. The prompt generation unit 106 may further reduce the data volume of prompt 303 by using high-compression lossless compression such as gzip. This allows the prompt generation unit 106 to significantly reduce the data volume of background image 307 through compression processing.

[0047] Fig. 4 is a diagram showing a flowchart of the restoration process of the first embodiment. The image restoration process of the first embodiment will be described with reference to Fig. 4. This flowchart is implemented using the background prompt generated by the flow shown in Fig. 2.

[0048] First, before the restoration process begins, a third trained model for generating a generated background image to be used by background generation unit 108 is stored in non-volatile memory 102. In addition, according to the flow shown in Figure 2, a main subject image, image size information of an input image, position information of the main subject image in the input image, and a prompt indicating background features are stored in non-volatile memory 102.

[0049] In S400, the background generation unit 108 generates a generated background image by restoring the background of the input image from the background image prompt. Specifically, the background generation unit 108 reads, from the non-volatile memory 102, a background image prompt indicating characteristics of the background of the input image and image size information of the input image. The background generation unit 108 reads a third trained model from the non-volatile memory 102 and inputs the prompt into the third trained model. As a result, the background generation unit 108 generates a generated background image according to the image size of the input image from the background image prompt by a single or a combination of multiple arithmetic operations. The generation of the generated background image will be described in detail later. The background generation unit 108 stores the generated background image in the memory 103.

[0050] Next, in S401, image restoration unit 109 restores an image similar to the input image from the main subject image and the generated background image to generate a restored image. Specifically, image restoration unit 109 reads the generated background image from memory 103. Image restoration unit 109 reads the main subject image and position information of the main subject in the input image from non-volatile memory 102. Image restoration unit 109 restores an image similar to the input image as a restored image by superimposing the main subject image on the generated background image at the same position as the position of the main subject in the input image indicated by the position information. This completes the restoration process. A detailed description of image restoration will be given later.

[0051] By this restoration process, image processing device 100 can restore an image close to the input image as a restored image from the main subject image, the data amount of which has been reduced compared to the data amount of the input image, and a prompt indicating the characteristics of the background.

[0052] 5 is a diagram illustrating the image processing steps of the restoration process of the first embodiment. The generation of a generated background image by prompt, which is the process of S400 in the flowchart of the restoration process, and the generation of a restored image by restoring an input image, which is the process of S401, will be described in detail with reference to FIG.

[0053] The image processing device 100 of this embodiment inputs the prompt 303 into a trained VLM (Visual Language Model) to generate noise data, and then generates and restores an image by having a trained diffusion model that uses the noise data as input perform de-diffusion processing. This is a typical example of Text to Image image generation. FIG. 5(a) shows a generated background image 500 generated by the background generation unit 108 using the prompt 303 as input. Note that the background generation unit 108 generates the generated background image 500 in which a background also exists in the area of ​​the main subject region image 305 that the main subject determination unit 104 cut out from the image of the background 302.

[0054] This process allows the background generation unit 108 to generate a generated background image 500 that is close to the background of the input image 300 from the prompt 303 with a small amount of data.

[0055] Next, image restoration unit 109 superimposes main subject image 306 shown in FIG. 3(c) on generated background image 500 shown in FIG. 5(a). Here, image restoration unit 109 superimposes main subject image 306 at the position of main subject 301 in input image 300 indicated by the position information of main subject image 306. In this way, image restoration unit 109 generates restored image 501 shown in FIG. 5(b), which is an image obtained by restoring the input image. Through this processing, image restoration unit 109 can generate restored image 501, an image that has few differences from input image 300.

[0056] As described above, according to the first embodiment, by prompting the background image of the input image, it is possible to suppress a reduction in the information of the background image, reduce the amount of data required to save the image, and restore an image that has few differences from the input image.

[0057] In the first embodiment, the main subject image is generated by performing transparency processing on the background and removing it, so the amount of data in the main subject image can be further reduced.

[0058] In the first embodiment, image size information of the input image and position information of the main subject image in the input image are saved, so the position of the main subject image and the image size of the restored image can be made closer to the input image.

[0059] In the first embodiment, the amount of output data is accepted from the user, so that it is possible to reduce the amount of data in accordance with the user's wishes.

[0060] In the first embodiment, if the data amount of the main subject image is not less than the output data amount, a new output data amount is obtained from the user. As a result, in the first embodiment, if the total data amount of the main subject image and the prompt data always exceeds the output data amount, the prompt processing can be omitted.

[0061] In the first embodiment, if the total data amount of the main subject image and the prompt is not less than the output data amount, the level of detail of the prompt is reduced and the prompt is generated again, thereby more reliably reducing the total data amount of the main subject image and the prompt.

[0062] Second Embodiment An image processing device according to a second embodiment will now be described. The second embodiment differs from the first embodiment in that the amount of data for prompts indicating background characteristics is automatically controlled based on the proportion of the separated main subject image and background image in the input image, rather than by a user instruction. The second embodiment also differs from the first embodiment in that the image restoration process does not restore the input image by superimposing the main subject on the generated background image, but instead extends the background remaining in the rectangularly cropped main subject region image to generate a background image. The hardware configuration of the second embodiment is the same as that of the first embodiment shown in FIG. 1(a).

[0063] FIG. 6 is a block diagram of the functional configuration of an image processing device according to the second embodiment. As shown in FIG. 6, the image processing device 100 according to the second embodiment includes a main subject determination unit 104, a prompt generation unit 106, and a background restoration unit 110. In other words, the image processing device 100 according to the second embodiment does not include the input unit 107, the background generation unit 108, and the image restoration unit 109 of the image processing device 100 according to the first embodiment. The main subject determination unit 104 according to the second embodiment performs different processing after separating the main subject region image from the background image. Therefore, the second embodiment will be described mainly with reference to the main subject determination unit 104 and the background restoration unit 110, and the description of the same content as in the first embodiment will be simplified. The background restoration unit 110 is an example of a restoration generation means.

[0064] In the second embodiment, main subject determination unit 104 separates an input image into a rectangular main subject region image including the main subject, and a background image including the background obtained by excluding the main subject region image from the input image. Main subject determination unit 104 compares the proportions of the separated main subject region image and background image in the input image, and determines the level of detail of the prompt for controlling the amount of data in the prompt. Main subject determination unit 104 temporarily stores the determined level of detail in memory 103.

[0065] The main subject determination unit 104 of the second embodiment does not perform transparency processing on the portions other than the main subject (i.e., the background portions) included in the rectangularly cut main subject region image, and stores the main subject region image in the non-volatile memory 102. In this embodiment, the main subject region image is an example of a first image.

[0066] The background restoration unit 110 acquires the third trained model stored in the non-volatile memory 102. The background restoration unit 110 acquires the main subject region image, image size information indicating the size of the input image, and a prompt, which is text data indicating the characteristics of the background, all stored in the non-volatile memory 102. The background restoration unit 110 inputs the prompt indicating the characteristics of the background in addition to the main subject region image into the third trained model and performs calculation processing. As a result, the background restoration unit 110 generates a restored image by restoring the input image by expanding the background remaining in the main subject region image to generate a background with the same image size as the input image. The background restoration unit 110 saves the generated restored image in the non-volatile memory 102.

[0067] 7 is a flowchart of image processing in the second embodiment. The processing from separating the main subject image and the background image to making the background prompt in the second embodiment will be described with reference to FIG.

[0068] Before image processing begins, a first trained model used by the main subject determination unit 104 to determine the input image and the main subject, and a second trained model used by the prompt generation unit 106 to generate a prompt indicating the characteristics of the background image based on the level of detail are stored in advance in the non-volatile memory 102.

[0069] In S700, main subject determination unit 104 separates a main subject region image and a background image included in the input image. Specifically, main subject determination unit 104 reads the input image and the first trained model from nonvolatile memory 102. Main subject determination unit 104 determines the main subject included in the input image by neural network arithmetic processing using the first trained model for the input image. Main subject determination unit 104 separates the main subject region image and the background image by cutting out the region in the input image where the main subject is located into a rectangle that fits the main subject. Thereafter, main subject determination unit 104 stores the generated main subject region image, image size information of the input image, and position information of the main subject or the main subject region image in the input image in nonvolatile memory 102. Main subject determination unit 104 also temporarily stores the background image in memory 103.

[0070] Next, in S701, main subject determination unit 104 determines the level of detail. Specifically, main subject determination unit 104 compares the proportions of the separated main subject region image and the background image in the input image, and determines the level of detail so that it becomes higher as the proportion of the background image increases. The determination of the level of detail will be described in detail later. Main subject determination unit 104 temporarily stores the determined level of detail in memory 103.

[0071] Next, in S702, the prompt generation unit 106 generates a prompt for the background image based on the level of detail. Specifically, the prompt generation unit 106 reads the background image and the level of detail from the memory 103. The prompt generation unit 106 reads the second trained model from the non-volatile memory 102. Using the second trained model that receives the background image as input, the prompt generation unit 106 generates a prompt, which is text data indicating features based on the level of detail, through neural network calculation processing. The prompt generation unit 106 stores the prompt indicating the features of the background in the non-volatile memory 102 and ends the flow.

[0072] The second trained model used by the prompt generation unit 106 generates a prompt indicating the characteristics of the background based on the level of detail, as in S203. The prompt generated by the prompt generation unit 106 has a larger amount of data as the level of detail increases. A detailed description of how the background image is converted into a prompt based on the level of detail will be given later.

[0073] By the above-described image processing flow, the image processing device 100 of the second embodiment can suppress the reduction of background information while reducing the total data amount of the main subject area image and the data amount of the prompt indicating the characteristics of the background to less than the data amount of the input image.

[0074] Fig. 8 is a diagram illustrating the image processing process and prompting in the image processing of the second embodiment. With reference to Fig. 8, the determination of the level of detail, which is the process of S701 in the flowchart, and the prompting of the background image based on the level of detail, which is the process of S702, will be described in detail.

[0075] Fig. 8(a) is an input image 800 to be processed in this embodiment. Fig. 8(b) is a main subject region image 805 including a main subject 801. Fig. 8(c) is a background image 806 including a background 802.

[0076] Input image 800 is an image taken during the daytime on a sunny day, with an apartment building and tree in the background and a dog on the grass in the foreground. The dog is located in the foreground and occupies a wide area of ​​the image. As a result of determining the main subject of input image 800, main subject determination unit 104 determines the dog located in the center as main subject 801. Main subject determination unit 104 cuts out a rectangular area from input image 800 that fits the dog, as shown in FIG. 8(b), to generate main subject region image 805 including main subject 801. Furthermore, main subject determination unit 104 generates a region from input image 800 excluding main subject region image 805 as background image 806 including background 802, as shown in FIG. 8(c). Main subject determination unit 104 compares the proportion of main subject region image 805 cut out in the rectangle shown in FIG. 8(b) with the proportion of background image 806 shown in FIG. 8(c) in the input image. Here, the proportion of the area image of main subject 801 is 40%, the proportion of the image of background 802 is 60%, and main subject determination unit 104 determines the level of detail to be 6. Note that main subject determination unit 104 may determine the level of detail using a table that associates the level of detail with the proportion, or a formula that calculates the level of detail from the proportion.

[0077] When determining the level of detail based on the input image of Fig. 3 in the first embodiment, main subject determination unit 104 determines the level of detail to be 9, assuming that the proportion of main subject region image 805 shown in Fig. 3(b) is 10% and the proportion of background image 806 shown in Fig. 3(d) is 90%. In this way, main subject determination unit 104 sets the level of detail to be higher the greater the proportion of background image in the input image.

[0078] Next, an example in which the prompt generation unit 106 generates a prompt indicating background characteristics from the background image 806 shown in Figure 8(c) based on the level of detail will be described with reference to Figures 8(d) and 8(e).

[0079] 8(d) is an example of a prompt 803 when the level of detail is low. When the level of detail is low, the prompt generation unit 106 generates a prompt in simple sentences that feature objects in the background, such as an apartment building, trees, and grass, and the positional relationships of simple objects. Therefore, the amount of information in the prompt that indicates the characteristics of the background is reduced, and the amount of data is also reduced.

[0080] 8(e) is an example of a prompt 804 when the level of detail is high. When the level of detail is high, the prompt generation unit 106 generates a detailed prompt. Specifically, the prompt generation unit 106 generates a detailed prompt indicating the characteristics of the background using features such as the background conditions (e.g., daytime and sunny weather), the number of apartment buildings (four in this case), the positions of the apartment buildings within the image, the relative positions of the apartment buildings, objects (e.g., trees and grass), the relative positions of the objects, the positional relationship between the apartment buildings and the objects, and other detailed relationships. Therefore, the amount of information in the prompt indicating the characteristics of the background increases, resulting in a large amount of data.

[0081] The image processing described above allows for a larger amount of information in the prompts that indicate the characteristics of the generated background, as the proportion of the background image in the input image increases, i.e., the degree of detail increases. Here, the larger the area of ​​the background that is not retained as an image, the larger the area of ​​the generated background image in the restored image. Therefore, in the second embodiment, when the background is large, detailed prompts with a large amount of information are retained, thereby reducing the difference between the restored image generated by the image restoration process and the input image.

[0082] Fig. 9 is a diagram showing a flowchart of restoration processing according to the second embodiment. The image restoration processing according to the second embodiment will be described using the flowchart in Fig. 9. The restoration processing is performed after the image processing shown in the flowchart in Fig. 7 is completed.

[0083] First, before the restoration process begins, a third trained model for expanding and generating a background to be used by background generation unit 108 is stored in advance in nonvolatile memory 102. In addition, according to the flow shown in Fig. 7, a main subject region image, image size information of the input image, position information of the main subject in the input image, and a prompt indicating the characteristics of the background are stored in nonvolatile memory 102.

[0084] In S900, background restoration unit 110 expands the background included in main subject region image 805 based on the background image prompt to generate a restored image by restoring the input image. Specifically, background restoration unit 110 reads from non-volatile memory 102 the main subject region image, image size information of the input image, position information of the main subject in the input image, and a background image prompt indicating background characteristics. Next, background restoration unit 110 inputs the main subject region image and the prompt into the third trained model read from non-volatile memory 102, thereby expanding the background remaining in the main subject region image through a single or a combination of multiple arithmetic operations to generate a background image. Through this expansion and generation of the background, a restored image is generated that has the same image size as the input image and in which the main subject is positioned at the position of the main subject in the input image, similar to the input image. A detailed description of the expansion and generation of the background image will be given later.

[0085] 10 is a diagram illustrating the image processing steps in the restoration process of the second embodiment. The expansion of the background image and the generation of the restored image in the restoration process of S900 in the flowchart will be described in detail with reference to FIG.

[0086] In this embodiment, a prompt 804 with a high level of detail is input to a trained VLM to generate noise data, and the noise data is added to the main subject region image 1000 and input to the trained diffusion model. The trained diffusion model then performs a de-diffusion process to generate an image that extends outside the main subject region image 1000. This is an example of image-to-image image generation.

[0087] FIG. 10(a) shows a main subject region image 1000 according to this embodiment. The main subject region image 1000 is cut out as a rectangle that fits the main subject 1001, and therefore includes a portion of the background 1002 that was also cut out. The background restoration unit 110 generates a restored image by expanding the background 1002 included in the main subject region image 1000 so that the image size of the input image and the position of the main subject in the input image are the same. FIG. 10(b) shows a restored image 1003 obtained by restoring the input image by expanding the background 1002 according to this embodiment. Increasing the amount of prompt information, as in this embodiment, makes the generated background 1002 more detailed and less blurred, thereby enabling the generation of a restored image 1003 that is closer to the input image. Therefore, even if the area of ​​the background not included in the main subject region image 1000 is large, increasing the amount of prompt information makes it possible to generate a restored image 1003 that differs little from the input image 800.

[0088] As described above, according to the second embodiment, the level of detail of the prompt indicating the characteristics of the background is set according to the proportion of the main subject and background image in the input image. As a result, the second embodiment can reduce the influence of the size of the main subject, reduce the loss of background information in the input image, significantly reduce the amount of data required to save the image, and generate a restored image by restoring an image with few differences from the input image.

[0089] In the second embodiment, the level of detail of the prompt indicating the characteristics of the background is set according to the ratio of the main subject and background image in the input image, thereby reducing the user's work such as setting the amount of output data.

[0090] In the second embodiment, transparency processing for removing the background from the main subject region image is omitted, so the processing load can be further reduced.

[0091] In the second embodiment, the input image is restored by expanding the background included in the main subject region image to generate a restored image, thereby reducing the difference from the input image, particularly around the main subject region image.

[0092] <Third embodiment> Next, an image processing device according to a third embodiment will be described with reference to Figs. 11 to 15. The third embodiment differs from the first and second embodiments in that an object in the background is recognized, and if a unique object is recognized, the name of the object is used to reduce the amount of information in the prompt. An example of the name of the object here is a proper noun. The hardware configuration of the second embodiment is the same as that of the first embodiment shown in Fig. 1(a).

[0093] 11 is a block diagram of the functional configuration of an image processing device of the third embodiment. Image processing device 100 of the third embodiment differs from image processing device 100 of the first embodiment in that input unit 107 is removed, and is otherwise similar to that of the first embodiment. However, the recognition processing in main subject determination unit 104 of the third embodiment and the generation processing of a prompt for a background image in prompt generation unit 106 are different from those of the first embodiment. Therefore, main subject determination unit 104 and prompt generation unit 106 will be described in detail, and the description of the other components will be simplified.

[0094] In the third embodiment, main subject determination unit 104 determines the main subject of an input image and separates the input image into a main subject image and a background image. After separation, main subject determination unit 104 recognizes objects appearing in the background image. Main subject determination unit 104 temporarily stores the recognition result, for example, the name of the object, in memory 103.

[0095] When the prompt generation unit 106 in the third embodiment recognizes a unique object that is uniquely determined in the background, it uses the name of the unique object to generate text data indicating the characteristics of the background as a prompt.

[0096] Fig. 12 is a diagram showing a flowchart of image processing in the third embodiment. Image processing in the third embodiment, from separating the main subject image from the background image to prompting the background, will be described using the flowchart in Fig. 12. Explanations of processes similar to those in the first embodiment will be simplified, and different processes will be described in detail.

[0097] First, a model is stored before image processing starts, as in the first embodiment. Also, step S1200 is the same as step S201 in the first embodiment.

[0098] Next, in S1201, main subject determination unit 104 recognizes an object appearing in the separated background image using image processing techniques such as pattern matching. Main subject determination unit 104 stores the name of the object and other information as the recognition result in memory 103. Recognition of an object appearing in the background image will be described in detail later.

[0099] Note that the recognition process for objects in the background image may use neural network calculation processing using a trained model other than pattern matching. In this case, the trained model is stored in the non-volatile memory 102 before the start of image processing.

[0100] Furthermore, in this embodiment, object recognition processing is performed only on the background image, but object recognition processing may also be performed on the main subject image in addition to the background image.

[0101] Next, in S1202, the prompt generation unit 106 first reads the recognition result, i.e., information on the names of objects appearing in the background image, from the memory 103, and determines whether or not there is a uniquely determined object based on the read object names. If the prompt generation unit 106 determines that there is no uniquely determined object, the process proceeds to S1203. On the other hand, if the prompt generation unit 106 determines that there is a uniquely determined object, the process proceeds to S1204.

[0102] The process of S1203 may be the same as the process of S203 in the first embodiment. In this embodiment, the setting of the level of detail is not particularly limited, but may be fixed at a medium setting, for example, rather than being automatically set based on a user input or a comparison of the sizes of the main subject image and the background image. Then, the prompt generation unit 106 stores text data indicating the characteristics of the generated background image as a prompt in the nonvolatile memory 102, and the flow ends.

[0103] In S1204, the prompt generation unit 106 generates a prompt for the background image using the uniquely determined name of the specific object. Specifically, the prompt generation unit 106 first reads the background image from the memory 103. Next, the prompt generation unit 106 reads the second trained model from the non-volatile memory 102, and generates text data indicating the characteristics of the background as a prompt using neural network arithmetic processing on the background image using the second trained model and the uniquely determined name of the specific object. Thereafter, the prompt generation unit 106 stores the generated prompt indicating the characteristics of the background in the non-volatile memory 102, and ends the flow.

[0104] A detailed description of prompting the background image using a unique object name will be given later.

[0105] Furthermore, when the prompt generation unit 106 performs object recognition processing on the main subject image in addition to the background image, if a uniquely determined object is captured in the main subject image, the prompt generation unit 106 may add to the text data indicating the characteristic that the background is that uniquely determined object.

[0106] 13 is a diagram illustrating the image processing process and prompts of the image processing of the third embodiment. The process of recognizing an object appearing in a background image, which is the process of S1201 in the flowchart, and the process of generating a prompt for the background image using a uniquely determined name of an object, which is the process of S1204, will be described in detail with reference to FIG.

[0107] FIG. 13(a) shows input image 1300 of this embodiment. Input image 1300 is an image taken during the daytime on a clear day, with Tokyo Tower in the background, trees on both sides, and a dog on the grass in front of Tokyo Tower. As a result of determining the main subject of input image 1300, main subject determination unit 104 determines the dog located in the foreground center as main subject 1301. Main subject determination unit 104 then cuts out the area of ​​input image 1300 where the dog is located into a rectangular area that fits the dog, as shown in FIG. 13(b), to generate main subject region image 1305. Then, of main subject region image 1305 cut out into the rectangle shown in FIG. 13(b), the portion other than main subject 1301, i.e., the small amount of background remaining, is subjected to transparency processing to generate main subject image 1306 shown in FIG. 13(c). The main subject determination unit 104 generates the remainder of the input image 1300 from which the main subject region image 1305 has been cut out as a background image 1307 including the background 1302 .

[0108] The processing up to this point is the same as in the first embodiment, and as a result, the main subject determination unit 104 can separate the input image 1300 shown in Figure 13(a) into the main subject image 1306 shown in Figure 13(c) and the background image 1307 shown in Figure 13(d).

[0109] Next, main subject determination unit 104 performs recognition processing using pattern matching on background image 1307 shown in FIG. 13(d). As a result, main subject determination unit 104 recognizes Tokyo Tower and a tree as objects. Here, Tokyo Tower is a uniquely determined object. Therefore, prompt generation unit 106 proceeds to processing to generate a prompt for the background using the name of the uniquely determined object.

[0110] The prompt generating unit 106 generates text data indicating the characteristics of the background shown in FIG. 13(e) as a prompt 1303 using the name of a specific object that is uniquely determined from the background image 1307 shown in FIG. 13(d).

[0111] As in prompt 1303 shown in FIG. 13(e), by using the name "Tokyo Tower," prompt generator 106 can omit or simplify text data indicating the tower's characteristics, such as its shape.

[0112] In this way, when an object included in the background image 1307 is a unique object, the image processing device 100 of the third embodiment generates a prompt 1303 including the name of the object, thereby reducing the amount of data in the prompt 1303, and thereby further reducing the amount of data to be saved.

[0113] Fig. 14 is a diagram showing a flowchart of restoration processing according to the third embodiment. Image restoration processing according to the third embodiment will be described using the flowchart in Fig. 14. This flowchart is premised on being performed after the flow shown in Fig. 12 has ended. The explanation of processing similar to that of the first embodiment will be simplified, and different processing will be described in detail.

[0114] First, before the start of the restoration process, a model is stored in advance, as in the third embodiment. Furthermore, the process of S1401 is the same as the process of S401 in the first embodiment.

[0115] Next, in S1400, the background generation unit 108 reads a prompt indicating the characteristics of the background and image size information of the input image from the non-volatile memory 102. Next, the background generation unit 108 reads the third trained model from the non-volatile memory 102 and inputs the prompt indicating the characteristics of the background, thereby generating a background image at the image size of the input image by a single or a combination of multiple arithmetic operations. Here, the prompt may be the name of an object included in the background image, and may include the name of a uniquely determined, specific object. In this case, the background generation unit 108 inputs the prompt including the name of the unique object into the third trained model, and generates a background image at the image size of the input image. A detailed description of the generation of a background image using a prompt including the name of a uniquely determined, specific object will be given later. The background generation unit 108 stores the generated background image in the memory 103.

[0116] Next, in S1401, the image restoration unit 109 restores the input image by the same process as in S401 in the first embodiment to generate a restored image, and then the flow ends.

[0117] 15 is a diagram illustrating the image processing steps of the restoration process of the third embodiment. Next, the generation of a background image by a prompt including a uniquely determined name of a specific object, which is the process of S1400 in the flowchart, will be described in detail with reference to FIG.

[0118] In this embodiment, as in the first embodiment, the prompt 1303 is input into a trained VLM (Visual Language Model) to generate noise data, and the trained diffusion model performs a de-diffusion process using the noise data as input to generate an image. The VLM and diffusion model of this embodiment have been trained on uniquely determined objects. FIG. 15(a) is a diagram of a generated background image 1500 generated using the prompt 1303 of this embodiment as input. In this embodiment, Tokyo Tower, which is a uniquely determined object, is included in the prompt, so a generated background image 1500 reproducing Tokyo Tower is generated.

[0119] This process allows the background generation unit 108 to generate a generated background image 1500 that is close to the background of the input image 1300 from a prompt 1303 with a smaller amount of data.

[0120] Figure 15(b) shows a restored image 1501 restored using processing similar to that of the first embodiment, and the image restoration unit 109 can restore an image with little difference from the input image 1300 using a generated background image 1500 generated from a prompt with a small amount of data.

[0121] As described above, according to the third embodiment, it is possible to further reduce the amount of prompt data while suppressing the reduction of background information in the input image, and to generate a restored image that has little difference from the input image.

[0122] <Fourth embodiment> An image processing device according to a fourth embodiment will now be described. The fourth embodiment differs from the first embodiment in that a main subject moving image and a background moving image are separated from each other and the background moving image is made into a prompt. The hardware configuration of the fourth embodiment is the same as that of the first embodiment shown in FIG. 1(a).

[0123] Fig. 16 is a block diagram of the functional configuration of an image processing device according to the fourth embodiment. As shown in Fig. 16, an image processing device 100 according to the fourth embodiment includes an input unit 107, a main subject determination unit 104, a prompt generation unit 106, a background generation unit 108, and an image restoration unit 109. The image processing device according to the fourth embodiment processes moving images. Therefore, in the fourth embodiment, the main subject determination unit 104, the prompt generation unit 106, the background generation unit 108, and the image restoration unit 109 will be mainly described, and the description of the same content as in the first embodiment will be simplified.

[0124] Main subject determination unit 104 determines the main subject in the input video by calculation using a first trained model. Based on the determination result, main subject determination unit 104 separates the input video into a main subject video including the main subject and a background video including the background, and generates a main subject video and a background video from the input video. The main subject video is an example of a first video, and the background video is an example of a second image. Main subject determination unit 104 stores the separated main subject video, video size information of the input video, and position information of the main subject in the input video in non-volatile memory 102. In this embodiment, main subject determination unit 104 stores each piece of information in non-volatile memory 102, but may also output and store it in an external storage device or a storage device on a network via output IF 111 or communication IF 112. In this embodiment, there are no particular limitations on the storage destination of the video size information of the input video, the position information of the main subject in the input video, and the main subject video. In addition, main subject determination unit 104 temporarily stores the background video in memory 103. The main subject determination unit 104 may determine the main subject using other moving image processing techniques such as pattern matching in addition to neural network calculation processing.

[0125] The prompt generation unit 106 generates text data indicating background characteristics in the background video as a prompt by computation using the second trained model. The prompt generation unit 106 saves the generated prompt indicating the background characteristics in the non-volatile memory 102. The prompt generation unit 106 saves the prompt in the non-volatile memory 102, but may also output and save the prompt to an external storage device or a network via the output IF 111 or the communication IF 112. Furthermore, the prompt generation unit 106 may include an expression indicating changes in the background over time to convert the video into a prompt. Furthermore, time information (time code) in the video may be recorded in the prompt, and multiple prompts corresponding to the time information may be generated. This is a prompt generation method for accurately expressing a background that changes over time. Furthermore, the prompt generation unit 106 may include information about the camera and lens used to capture the video. This information may be read from the camera settings at the time of capture, from metadata saved in the video, or estimated from the content of the captured video; the means are not limited. Furthermore, the prompt generation unit 106 may include camerawork information in the prompt. In this embodiment, there are no particular limitations on where the generated prompt indicating the characteristics of the background is saved. The prompt generation unit 106 may generate the prompt indicating the characteristics of the background using other video processing techniques, such as pattern matching, in addition to neural network arithmetic processing.

[0126] The background generation unit 108 generates a generated background video that restores the background by inputting a prompt, which is text data indicating the characteristics of the background, into a third trained model and performing arithmetic processing. The background generation unit 108 temporarily stores the generated background video in the memory 103. The third trained model may be a single trained model or a combination of multiple trained models such as a VLM (Visual Language Model) and a diffusion model. An example of the third trained model may be Sora, developed by OpenAI (registered trademark).

[0127] Image restoration unit 109 generates a restored moving image by restoring the input moving image using the generated background moving image and the main subject moving image. Specifically, image restoration unit 109 generates the restored moving image by superimposing the main subject moving image on the generated background moving image at a position corresponding to the position information of the main subject in the input moving image. Image restoration unit 109 stores the generated restored moving image in non-volatile memory 102. Image restoration unit 109 may output the generated restored moving image to an external display device or the like via output IF 111 for display.

[0128] Next, image processing from separating the main subject and background from an input video to prompting the background in the fourth embodiment will be described. Fig. 17 is a flowchart of the image processing in the fourth embodiment.

[0129] Before image processing begins, the input video to be used by the main subject determination unit 104, a first trained model for determining the main subject, and a second trained model for generating text data indicating the characteristics of the background video to be used by the prompt generation unit 106 are pre-stored in the non-volatile memory 102.

[0130] In S1700, input unit 107 acquires from the user an output data amount as a threshold for the total data amount of the main subject moving image data and the text data amount indicating the characteristics of the background moving image. Input unit 107 stores the acquired output data amount in memory 103.

[0131] Next, in S1701, main subject determination unit 104 determines a main subject in the input video, separates the main subject video from the background video, and generates a main subject video and a background video from the input video. Specifically, main subject determination unit 104 reads the input video and a first trained model from non-volatile memory 102. Main subject determination unit 104 uses the first trained model on the input video to determine the main subject in the input video through neural network calculation processing. Main subject determination unit 104 removes the background from the main subject region video, in other words, performs transparency processing, to generate the main subject video. Main subject determination unit 104 stores the generated main subject video and video size information of the input video in non-volatile memory 102. Main subject determination unit 104 also temporarily stores the background video in memory 103. A detailed description of the separation of the main subject video from the background video will be given later.

[0132] Next, in S1702, main subject determination unit 104 determines whether the data amount of the main subject moving image is less than the output data amount. Specifically, main subject determination unit 104 reads the output data amount received from the user from memory 103. Main subject determination unit 104 compares the data amount of the generated main subject moving image with the output data amount. If main subject determination unit 104 determines that the data amount of the main subject moving image is not less than the output data amount, the process returns to S1700 and prompts the user to input the output data amount again. This causes input unit 107 to acquire a new output data amount. Note that if steps S1700 and subsequent steps are repeated, main subject determination unit 104 may omit S1701. On the other hand, if main subject determination unit 104 determines that the data amount of the main subject moving image is less than the output data amount, the process proceeds to S1703.

[0133] Next, in S1703, the prompt generation unit 106 converts the background video into a prompt. Specifically, the prompt generation unit 106 reads the background video from the memory 103 and also reads the second trained model from the non-volatile memory 102. The prompt generation unit 106 uses the second trained model for the background video to generate a prompt in which the features of the background video are converted into text through neural network calculation processing. Therefore, the prompt can also be considered text data. The prompt generation unit 106 stores the generated prompt indicating the features of the background video in the non-volatile memory 102. Here, the second trained model may generate a prompt indicating the features of the background based on the level of detail. The user may set the level of detail via the input unit 107. The higher the level of detail, the greater the amount of data in the prompt indicating the features of the background. A detailed description of converting the background video into a prompt will be given later.

[0134] Next, in S1704, prompt generation unit 106 determines whether the total data amount of the main subject moving image and the prompt is less than the output data amount. Specifically, prompt generation unit 106 reads the output data amount received from the user from memory 103. Prompt generation unit 106 calculates the sum of the data amount of the main subject moving image read from non-volatile memory 102 and the data amount of the generated prompt indicating the characteristics of the background. Prompt generation unit 106 compares the calculated total data amount with the output data amount. If prompt generation unit 106 determines that the calculated total data amount is not less than the output data amount, the process proceeds to S205.

[0135] In S1705, the prompt generation unit 106 reduces the level of detail of the prompt, and the process proceeds to S1703. The prompt generation unit 106 may reduce the level of detail based on a preset setting value. Alternatively, the prompt generation unit 106 may notify the user to reduce the level of detail, and reduce the level of detail based on the level of detail input by the user.

[0136] In S1703, the prompt generation unit 106 generates a prompt indicating the characteristics of the background by converting the background video into a prompt based on the lowered level of detail. Here, since the prompt generated by the prompt generation unit 106 has a low level of detail, the data amount of the prompt is small.

[0137] Thereafter, prompt generating unit 106 repeats S1705 and S1703 until the total data amount of the main subject moving image and the prompt becomes smaller than the output data amount.

[0138] In S1704, if prompt generating unit 106 determines that the total data amount of the main subject moving image and the prompt is less than the output data amount, the flow ends.

[0139] As a result, the image processing device 100 reduces the total data amount of the main subject video and the prompt data amount indicating the background characteristics from the output data amount received from the user, thereby reducing the data amount of the input video while suppressing a reduction in background information.

[0140] 18 is a diagram illustrating the moving image processing process and prompting in the image processing of the fourth embodiment. Next, the separation of the main subject moving image and the background moving image, which is the process in S1701 of the flowchart, and the prompting of the background moving image, which is the process in S1703, will be described in detail with reference to FIG.

[0141] FIG. 18(a) shows an input video of this embodiment. Example frames of the video are arranged in chronological order as 1800, 1810, and 1820. These frames are representative frames, and there may be several frames in between. The input video shows an apartment building in the background, with scattered trees in front of the apartment building, and a dog walking on the lawn in front of the trees. Assume that main subject determination unit 104 determines the main subjects of input video frames 1800, 1810, and 1820, and as a result, determines the dog located in the center as main subject 1802, 1812, and 1822.

[0142] Of the input video frames 1800, 1810, and 1820 containing the main subjects shown in Fig. 18(a), main subject determination unit 104 performs transparency processing on the portions other than main subjects 1802, 1812, and 1822, i.e., the background portions. As a result, main subject determination unit 104 generates main subject video frames 1830, 1840, and 1850, which are videos of main subjects 1802, 1812, and 1822 shown in Fig. 18(b).

[0143] Main subject determination section 104 generates the remaining video, after cutting out main subjects 1802, 1812, and 1822 from input video frames 1800, 1810, and 1820, as background video frames 1860, 1870, and 1880, which are video of backgrounds 1801, 1811, and 1821 shown in FIG. 18(c).

[0144] As a result, the main subject determination unit 104 separates the input video frames 1800, 1810, and 1820 shown in Figure 18(a) into main subject video frames 1830, 1840, and 1850, which are videos of the main subjects 1802, 1812, and 1822 shown in Figure 18(b), and background video frames 1860, 1870, and 1880, which are videos of the backgrounds 1801, 1811, and 1821 shown in Figure 18(c).

[0145] Next, the prompt generating unit 106 generates text data indicating the characteristics from the moving images of the backgrounds 1801, 1811, and 1821 shown in FIG. 18(c) as prompt 1890 shown in FIG. 18(d).

[0146] As shown in FIG. 18(d), the prompt generation unit 106 generates prompt 1890, which is text data, based on background video frames 1860, 1870, and 1880. The prompt generation unit 106 includes features indicating the background situation, features of objects present in the background, the relative positions of the objects, and time-series changes in the background. Specifically, based on background video frames 1860, 1870, and 1880 of FIG. 18(c), the prompt generation unit 106 generates features such as four apartment buildings in the background, with the left building at the back and the right building at the foreground, scattered trees in front of the apartment buildings, and grass in front of those. The prompt generation unit 106 also generates prompt 1890 with features such as the camera panning from right to left and Tokyo Tower entering the frame from the left in background video frame 1870. Because prompt 1890 is text data, the data volume is very small. The prompt generation unit 106 may further reduce the data volume of prompt 303 by using high-compression lossless compression such as gzip. This allows prompt generation unit 106 to significantly reduce the amount of data in background video frames 1860, 1870, and 1880 through compression processing. One or more frames of the background video may be saved as images in memory 103. These may be used as auxiliary information when generating a background video from a prompt, as described below. The timing for saving one or more frames is not limited and may be at a predetermined time interval or when the background video changes significantly.

[0147] Fig. 19 is a diagram showing a flowchart of the restoration process of the fourth embodiment. The restoration process of a video in the fourth embodiment will be described with reference to Fig. 19. This flowchart is implemented using the background prompt generated by the flow shown in Fig. 17.

[0148] First, before the restoration process begins, a third trained model for generating a generated background video to be used by background generation unit 108 is stored in non-volatile memory 102. In addition, according to the flow shown in Figure 17, a main subject video, video size information of the input video, and a prompt indicating the characteristics of the background are stored in non-volatile memory 102.

[0149] In S1900, the background generation unit 108 generates a generated background video by restoring the background of the input video from the background video prompt. Specifically, the background generation unit 108 reads, from the non-volatile memory 102, a background video prompt indicating background characteristics of the input video and video size information of the input video. The background generation unit 108 reads the third trained model from the non-volatile memory 102 and inputs the prompt into the third trained model. As a result, the background generation unit 108 generates a generated background video according to the video size of the input video from the background video prompt by a single or a combination of multiple arithmetic operations. A detailed description of the generation of the generated background video will be given later. The background generation unit 108 stores the generated background video in the memory 103. Here, if the third trained model supports image-based video generation (Image to Video), one or more frame images of the above-mentioned background video may be input as auxiliary information to generate the generated background video together with the prompt.

[0150] Next, in S1901, image restoration unit 109 restores a moving image similar to the input moving image from the main subject moving image and the generated background moving image to generate a restored moving image. Specifically, image restoration unit 109 reads the generated background moving image from memory 103. Image restoration unit 109 reads the main subject moving image from non-volatile memory 102. Image restoration unit 109 restores a moving image similar to the input moving image as a restored moving image by superimposing the main subject moving image on the generated background moving image. This completes the restoration process. A detailed description of moving image restoration will be given later.

[0151] Through this restoration process, image processing device 100 can restore a restored video that is close to the input video from the main subject video, the data amount of which has been reduced compared to the data amount of the input video, and the prompt indicating the characteristics of the background.

[0152] Fig. 20 is a diagram illustrating the moving image processing steps of the restoration processing of the fourth embodiment. The generation of a generated background moving image by a prompt, which is the processing of S1900 in the flowchart of the restoration processing, and the generation of a restored moving image by restoring an input moving image, which is the processing of S1901, will be described in detail with reference to Fig. 20.

[0153] The image processing device 100 of this embodiment inputs a prompt 303 into a trained VLM (Visual Language Model) to generate noise data, and then generates and restores a moving image by having a trained diffusion model that uses the noise data as input perform a de-diffusion process. This is a typical example of Text to Video moving image generation. FIG. 20(a) shows generated background moving image frames 2000, 2010, and 2020 that the background generation unit 108 generated using the prompt 303 as input. Note that the background generation unit 108 generates generated background moving image frames 2000, 2010, and 2020 in which background also exists in the areas of main subjects 1802, 1812, and 1822 that the main subject determination unit 104 cut out from the moving images of backgrounds 1801, 1811, and 1821.

[0154] This process allows the background generation unit 108 to generate background video frames 1860, 1870, and 1880 that are close to the backgrounds of the input video frames 1800, 1810, and 1820 from a prompt 1890 with a small amount of data.

[0155] Next, image restoration unit 109 superimposes main subject video frames 1830, 1840, and 1850 shown in Fig. 18(b) on background video frames 1860, 1870, and 1880 shown in Fig. 20(a). In this way, image restoration unit 109 generates restored video frames 2030, 2040, and 2050 shown in Fig. 20(b), which are video obtained by restoring the input video. Through this processing, image restoration unit 109 can generate restored video frames 2030, 2040, and 2050, which are video that differs little from input video frames 1800, 1810, and 1820.

[0156] As described above, according to the fourth embodiment, by prompting the background video of the input video, it is possible to suppress the reduction in information of the background video, reduce the amount of data required to store the video, and restore a video that has little difference from the input video.

[0157] In the fourth embodiment, a main subject moving image is generated in which the background is transparently removed, so that the data amount of the main subject moving image can be further reduced.

[0158] In the fourth embodiment, the amount of output data is accepted from the user, so that it is possible to reduce the amount of data in accordance with the user's wishes.

[0159] In the fourth embodiment, if the data amount of the main subject moving image is not less than the output data amount, a new output data amount is obtained from the user. As a result, in the fourth embodiment, if the total data amount of the main subject moving image and the prompt data always exceeds the output data amount, the prompt processing can be omitted.

[0160] In the fourth embodiment, if the total data volume of the main subject video and the prompt is not less than the output data volume, the level of detail of the prompt is reduced and the prompt is generated again, thereby more reliably reducing the total data volume of the main subject video and the prompt.

[0161] Fifth Embodiment An image processing device according to a fifth embodiment will be described. The fifth embodiment differs from the fourth embodiment in that, in addition to the fourth embodiment, audio data accompanying the input video is used to improve the accuracy of the generated background video. The hardware configuration of the fifth embodiment is the same as that of the first embodiment shown in FIG. 1(a).

[0162] FIG. 21 is a block diagram of the functional configuration of an image processing device according to a fifth embodiment. As shown in FIG. 21, an image processing device 100 according to the fifth embodiment includes an input unit 107, a main subject determination unit 104, a prompt generation unit 106, a background generation unit 108, an image restoration unit 109, and an audio input unit 2100. The image processing device according to the fifth embodiment uses audio data in addition to processing moving images. Therefore, the fifth embodiment will be described mainly with reference to the prompt generation unit 106, the background generation unit 108, the image restoration unit 109, and the audio input unit 2100, and the description of the same content as in the fourth embodiment will be simplified.

[0163] The audio input unit 2100 simultaneously records audio when recording video. The audio input device may be a microphone built into the image processing device or a microphone attached externally to the audio recording device, and is not limited thereto. The audio input by the audio input unit 2100 is stored as various pieces of information in the nonvolatile memory 102, but may also be output and stored in an external storage device or a storage device on a network via the output IF 111 or the communication IF 112.

[0164] As in the fourth embodiment, the prompt generation unit 106 generates text data indicating background features in the background video as a prompt by computational processing using the second trained model. In the fifth embodiment, the prompt generation unit 106 may save the audio prompt, which is the input audio prompt, and the background video prompt, which is the background video prompt, as separate files, or may edit and save the prompt as a single file. Specifically, for background video prompts, if the audio includes birdsong, text describing environmental sounds, such as birdsong, can be added to the prompt. Alternatively, if the audio includes human speech, information that a person is speaking and the content of the speech can be converted into text and added directly to the prompt. Alternatively, if a sound is generated at a specific timing, both time information, such as the time code of the input video, and text describing the generated sound can be added to the prompt. For example, the prompt generation unit 106 may generate prompts indicating audio features using other audio processing techniques, such as spectral analysis, in addition to neural network computational processing.

[0165] The background generation unit 108 generates a generated background video with the background restored by inputting a prompt, which is text data indicating the characteristics of the background, into a third trained model and performing arithmetic processing. The background generation unit 108 in the fifth embodiment can further input a prompt including the above-mentioned audio information to generate a generated background video with the background restored. Alternatively, the background generation unit 108 can input the input audio as is to generate a generated background video. When generating a background video by adding audio information, the background generation unit 108 can estimate the environment in which the video was shot from text describing environmental sounds or the environmental sounds themselves, and use this as auxiliary information for generating the background video. Furthermore, the background generation unit 108 can estimate the relationship between people appearing in the background and their origin from text describing people's voices or the voices themselves, and use this as auxiliary information for generating the mouth movements and facial expressions of people appearing in the background. Furthermore, the background change timing can be estimated from text describing the time information of the audio or the audio itself, and use this as auxiliary information for generating the changing background video.

[0166] The image restoration unit 109 generates a restored video by restoring the input video using the generated background video and the main subject video. In the fifth embodiment, the image restoration unit 109 also adds audio to the restored video. The input audio stored in the non-volatile memory 102 or the like may be added directly to the restored video. Furthermore, if the image processing device is equipped with an audio restoration unit (not shown) that can restore audio from text using neural network arithmetic processing or the like, audio may be restored from an audio prompt generated by the prompt generation unit 106 and the restored audio may be added to the restored video. The method of adding audio to the restored video is not limited, and may be added using a video format that can store audio, such as MPEG4, or by associating the video and audio using other metadata or the like.

[0167] Next, the image processing in the fifth embodiment, from separating the main subject and background from the input video to prompting the background and audio, will be described. Fig. 22 is a diagram of a flowchart of the image processing in the fifth embodiment. Also, explanations of steps similar to those in Fig. 17, which is the flowchart of the fourth embodiment, will be omitted. Furthermore, if audio data is not made into a prompt and audio data is directly input when restoring the background video, the audio data is not made into a prompt, so the flowchart in Fig. 17 may be applied instead of the flowchart in Fig. 22.

[0168] Before image processing begins, the input video to be used by the main subject determination unit 104, a first trained model for determining the main subject, and a second trained model for generating text data indicating the characteristics of the background video and audio data to be used by the prompt generation unit 106 are pre-stored in the non-volatile memory 102.

[0169] Steps S1700 to S1703 are the same as those in FIG. 22 of the fourth embodiment, and therefore the description thereof will be omitted.

[0170] Next, in S2200, the prompt generation unit 106 converts the voice data into a prompt. Specifically, the prompt generation unit 106 reads the voice data from the memory 103 and also reads the second trained model from the non-volatile memory 102. The prompt generation unit 106 uses the second trained model on the voice data to generate a prompt in which the voice features are converted into text through neural network calculation processing. Therefore, the prompt can also be considered text data. The prompt generation unit 106 stores a prompt indicating the features of the voice data in the non-volatile memory 102. Here, the second trained model may generate a prompt indicating the voice features based on the level of detail. The user may set the level of detail via the input unit 107. The higher the level of detail, the greater the amount of data in the prompt indicating the voice features. A detailed description of converting voice into a prompt will be given later.

[0171] Next, in S1704, prompt generation unit 106 determines whether the total data amount of the main subject moving image and the prompt is less than the output data amount. Specifically, prompt generation unit 106 reads the output data amount received from the user from memory 103. Prompt generation unit 106 calculates the sum of the data amount of the main subject moving image read from non-volatile memory 102 and the data amount of the generated prompt indicating the background and audio features. Prompt generation unit 106 compares the calculated total data amount with the output data amount. If prompt generation unit 106 determines that the calculated total data amount is not less than the output data amount, the process proceeds to S1705.

[0172] In S1705, the prompt generation unit 106 reduces the level of detail of the prompt, and the process proceeds to S1703. The prompt generation unit 106 may reduce the level of detail based on a preset setting value. Alternatively, the prompt generation unit 106 may notify the user to reduce the level of detail, and reduce the level of detail based on the level of detail input by the user.

[0173] In S1703, the prompt generation unit 106 converts the background video and audio data into a prompt based on the lowered level of detail, and re-generates a prompt that indicates the characteristics of the background and audio. Here, since the prompt generated by the prompt generation unit 106 has a low level of detail, the data amount of the prompt is small.

[0174] Thereafter, prompt generating unit 106 repeats S1705, S1703, and S2200 until the total data amount of the main subject moving image and the prompt becomes smaller than the output data amount.

[0175] In S1704, if prompt generating unit 106 determines that the total data amount of the main subject moving image and the prompt is less than the output data amount, the flow ends.

[0176] As a result, the image processing device 100 reduces the total data amount of the main subject video and the data amount of prompts indicating background and audio characteristics from the amount of output data received from the user, thereby reducing the data amount of the input video while suppressing a reduction in background and audio information.

[0177] Fig. 23 is a diagram showing a flowchart of the restoration process of the fifth embodiment. The restoration process of a video in the fifth embodiment will be described with reference to Fig. 23. This flowchart is implemented using the background and audio prompts generated by the flow shown in Fig. 22.

[0178] First, before the restoration process begins, a third trained model for generating a generated background video to be used by the background generation unit 108 is stored in advance in the non-volatile memory 102. In addition, according to the flow shown in Figure 22, the main subject video, video size information of the input video, and prompts indicating the characteristics of the background and audio are stored in the non-volatile memory 102.

[0179] In S2300, the background generation unit 108 generates a generated background video by restoring the background of the input video from the prompt and audio data of the background video. Alternatively, the background generation unit 108 may use a prompt indicating the audio characteristics generated in FIG. 22 instead of using the audio data itself. Specifically, the background generation unit 108 reads a background video prompt indicating the background characteristics of the input video, audio data or an audio prompt, and video size information of the input video from the non-volatile memory 102. The background generation unit 108 reads a third trained model from the non-volatile memory 102 and inputs the prompt and, if audio data itself is used, the audio data into the third trained model. As a result, the background generation unit 108 generates a generated background video according to the video size of the input video from the prompt of the background video and the audio data or an audio prompt by a single or a combination of multiple arithmetic processes. A detailed description of the generation of the generated background video will be given later. The background generation unit 108 stores the generated background video in the memory 103. Here, if the third trained model supports image-based video generation (Image to Video), one or more frame images of the aforementioned background video may be input as auxiliary information to generate a generated background video together with the prompt.

[0180] Next, in S2301, image restoration unit 109 restores a video similar to the input video from the main subject video, the generated background video, and the audio to generate a restored video. Specifically, image restoration unit 109 reads the generated background video from memory 103. Image restoration unit 109 reads the main subject video from non-volatile memory 102. Image restoration unit 109 superimposes the main subject video on the generated background video to restore a video similar to the input video as a restored video. Image restoration unit 109 also reads audio data from memory 103 and adds it to the restored video. This completes the restoration process.

[0181] Through this restoration process, the image processing device 100 can restore a restored video that is closer to the input video from the main subject video, which has a reduced data volume compared to the data volume of the input video, a prompt indicating the characteristics of the background, and audio.

[0182] As described above, according to the fifth embodiment, by further using the features of audio data in addition to the fourth embodiment, it is possible to restore a moving image that has fewer differences from the input moving image.

[0183] (Other embodiments) The above-described embodiments may be combined. For example, the first to fifth embodiments may be combined so that the user can select one of the methods of each embodiment.

[0184] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. The present disclosure can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0185] The disclosure of this specification includes the following image processing device, image processing method, and program. (Item 1) a determining means for determining a main subject of an input image including a main subject and a background, and generating a first image including the main subject and a second image including the background from the input image; generating means for generating text data indicating characteristics of the background as a prompt based on the second image; An image processing device comprising: (Item 2) The determining means generates an image obtained by removing a background from an image of a rectangular area including the main subject as the first image. 2. The image processing device according to item 1, (Item 3) The determining means generates an image of a rectangular area including the main subject as the first image. 3. The image processing device according to item 1 or 2, (Item 4) The determining means stores image size information of the input image and position information of the main subject in the input image. 4. The image processing device according to any one of items 1 to 3, wherein: (Item 5) a restoration generating means for generating an image of the background as a third image based on the prompt, and superimposing the first image on the third image to generate a restored image obtained by restoring the input image; 3. The image processing device according to item 2, comprising: (Item 6) the determining means outputs image size information of the input image and position information of the main subject in the input image; The restoration generating means generates the third image based on the image size information, and superimposes the first image on the third image based on the position information. 6. The image processing device according to item 5, (Item 7) a restoration generating means for expanding a part of the background included in the first image based on the prompt and generating a restored image by restoring the input image; 4. The image processing device according to item 3, comprising: (Item 8) the determining means outputs image size information of the input image and position information of the main subject in the input image; The restoration generating means extends a part of the background included in the first image based on the image size information and the position information. 8. The image processing device according to item 7, (Item 9) an input means for acquiring an amount of output data input by a user; The generating means generates the prompt based on the amount of output data. 9. The image processing device according to any one of items 1 to 8, wherein: (Item 10) The generating means generates the prompt so that the total data amount of the first image and the prompt does not exceed the output data amount. 10. The image processing device according to item 9, (Item 11) The generating means controls the amount of data of the prompt to be generated in accordance with the ratio of the sizes of the first image and the second image in the input image. 11. The image processing device according to any one of items 1 to 10, wherein: (Item 12) the determining means recognizes an object included in the background and stores the name of the object; The generating means generates the prompt based on the name of the object. 2. The image processing device according to item 1, (Item 13) When the name of the object is a unique name that uniquely defines the object, the generating means generates the prompt based on the name. Item 13. The image processing device according to item 12. (Item 14) When the data amount of the first image is not less than the output data amount, the input means acquires a new output data amount from the user. 10. The image processing device according to item 9, (Item 15) When the amount of data of the first image and the second image is not less than the amount of output data, the generating means reduces the level of detail of the prompt and generates the prompt again. 10. The image processing device according to item 9, (Item 16) the input image includes a moving image; the determining means determines a main subject of an input moving image including a main subject and a background, and generates a first moving image including the main subject and a second moving image including the background from the input moving image; The generating means generates text data indicating characteristics of the background as a prompt based on the second video. 2. The image processing device according to item 1, (Item 17) The determining means generates a moving image by deleting a background from a moving image of a rectangular area including the main subject as the first moving image. Item 17. The image processing device according to item 16, (Item 18) The determining means generates a moving image of a rectangular area including the main subject as the first moving image. 18. The image processing device according to item 16 or 17, (Item 19) The determining means stores information about the size of the input moving image and information about the position of the main subject in the input moving image. 19. The image processing device according to any one of items 16 to 18, wherein: (Item 20) a restoration generating means for generating the background moving image as a third moving image based on the prompt, and superimposing the first moving image on the third moving image to generate a restored moving image by restoring the input moving image; Item 18. The image processing device according to item 17, comprising: (Item 21) the determining means outputs moving image size information of the input moving image and position information of the main subject in the input moving image; The restoration generating means generates the third moving image based on the moving image size information, and superimposes the first moving image on the third moving image based on the position information. 21. The image processing device according to item 20, characterized in that (Item 22) The generating means includes generating means for generating text data indicating characteristics of the audio as a prompt from audio data accompanying the input video. 22. The image processing device according to any one of items 16 to 21, wherein: (Item 23) a restoration generating means for generating the background moving image as a third moving image based on a prompt indicating a feature of the background and audio data accompanying the input moving image, restoring the input moving image by superimposing the first moving image on the third moving image, and adding the audio data to generate the restored moving image; 23. The image processing device according to any one of items 16 to 22, comprising: (Item 24) a restoration generating means for generating the background moving image as a third moving image based on a prompt indicating a characteristic of the background and a prompt indicating a characteristic of audio generated from audio data accompanying the input moving image, restoring the input moving image by superimposing the first moving image on the third moving image, and adding the audio data accompanying the input moving image to generate the restored moving image; 23. The image processing device according to any one of items 16 to 22, comprising: (Item 25) a determining step of determining a main subject of an input image including a main subject and a background, and generating a first image including the main subject and a second image including the background from the input image; a generating step of generating text data indicating characteristics of the background as a prompt based on the second image; An image processing method comprising: (Item 26) A program for causing a computer to function as each means of the image processing device described in item 1.

[0186] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0187] 100... Image processing device, 101... Processor, 104... Main subject determination unit, 106... Prompt generation unit, 107... Input unit, 108... Background generation unit, 109... Image restoration unit, 110... Background restoration unit, 300, 800, 1300... Input image, 301, 801, 1001, 1301, 1802, 1812, 1822... Main subject, 302, 802, 1002, 1302, 1801, 1811, 1821... Background, 303, 803, 804, 1303, 1890... Prompt, 305, 805, 1000, 1305, 1830, 1840, 1850...Main subject region image, 306, 1306...Main subject image, 307, 806, 1307...Background image, 500, 1500...Generated background image, 501, 1003, 1501...Restored image, 1800, 1810, 1820...Input video frame, 1860, 1870, 1880...Background video frame, 2000, 2010, 2020...Generated background video frame, 2030, 2040, 2050...Restored video frame, 2100...Audio input unit.

Claims

1. a determining means for determining a main subject of an input image including a main subject and a background, and generating a first image including the main subject and a second image including the background from the input image; a generating means for generating text data indicating characteristics of the background as a prompt based on the second image; An image processing device comprising:

2. The determining means generates an image obtained by removing a background from an image of a rectangular area including the main subject as the first image.

2. The image processing device according to claim 1, wherein:

3. The determining means generates an image of a rectangular area including the main subject as the first image.

2. The image processing device according to claim 1, wherein:

4. The determining means stores image size information of the input image and position information of the main subject in the input image.

2. The image processing device according to claim 1, wherein:

5. a restoration generating means for generating an image of the background as a third image based on the prompt, and superimposing the first image on the third image to generate a restored image obtained by restoring the input image; 3. The image processing device according to claim 2, further comprising:

6. the determining means outputs image size information of the input image and position information of the main subject in the input image; The restoration generating means generates the third image based on the image size information, and superimposes the first image on the third image based on the position information.

6. The image processing device according to claim 5,

7. a restoration generating means for expanding a portion of the background included in the first image based on the prompt and generating a restored image by restoring the input image; 4. The image processing device according to claim 3, further comprising:

8. the determining means outputs image size information of the input image and position information of the main subject in the input image; The restoration generating means extends a part of the background included in the first image based on the image size information and the position information.

8. The image processing device according to claim 7,

9. an input means for acquiring an amount of output data input by a user; The generating means generates the prompt based on the amount of output data.

2. The image processing device according to claim 1, wherein:

10. The generating means generates the prompt so that the total data amount of the first image and the prompt does not exceed the output data amount.

10. The image processing device according to claim 9,

11. The generating means controls the amount of data of the prompt to be generated in accordance with the ratio of the sizes of the first image and the second image in the input image.

2. The image processing device according to claim 1, wherein:

12. the determining means recognizes an object included in the background and stores the name of the object; The generating means generates the prompt based on the name of the object.

2. The image processing device according to claim 1, wherein:

13. When the name of the object is a unique name that uniquely defines the object, the generating means generates the prompt based on the name.

13. The image processing device according to claim 12.

14. When the data amount of the first image is not less than the output data amount, the input means acquires a new output data amount from the user.

10. The image processing device according to claim 9,

15. If the amount of data of the first image and the second image is not less than the amount of output data, the generating means reduces the level of detail of the prompt and generates the prompt again.

10. The image processing device according to claim 9,

16. the input image includes a moving image; the determining means determines a main subject of an input moving image including a main subject and a background, and generates a first moving image including the main subject and a second moving image including the background from the input moving image; The generating means generates text data indicating characteristics of the background as a prompt based on the second moving image.

2. The image processing device according to claim 1, wherein:

17. The determining means generates a moving image by deleting a background from a moving image of a rectangular area including the main subject as the first moving image.

17. The image processing device according to claim 16,

18. The determining means generates a moving image of a rectangular area including the main subject as the first moving image.

17. The image processing device according to claim 16,

19. The determining means stores information about the size of the input moving image and information about the position of the main subject in the input moving image.

17. The image processing device according to claim 16,

20. a restoration generating means for generating the background moving image as a third moving image based on the prompt, and superimposing the first moving image on the third moving image to generate a restored moving image by restoring the input moving image; 18. The image processing device according to claim 17, further comprising:

21. the determining means outputs moving image size information of the input moving image and position information of the main subject in the input moving image; The restoration generating means generates the third moving image based on the moving image size information, and superimposes the first moving image on the third moving image based on the position information.

21. The image processing device according to claim 20.

22. The generating means includes generating means for generating text data indicating characteristics of the audio as a prompt from audio data accompanying the input video.

17. The image processing device according to claim 16,

23. a restoration generating means for generating the background moving image as a third moving image based on a prompt indicating a feature of the background and audio data accompanying the input moving image, restoring the input moving image by superimposing the first moving image on the third moving image, and adding the audio data to generate the restored moving image; 23. The image processing device according to claim 16, further comprising:

24. a restoration generating means for generating the background moving image as a third moving image based on a prompt indicating a characteristic of the background and a prompt indicating a characteristic of an audio generated from audio data accompanying the input moving image, restoring the input moving image by superimposing the first moving image on the third moving image, and adding the audio data accompanying the input moving image to generate the restored moving image; 23. The image processing device according to claim 16, further comprising:

25. a determining step of determining a main subject of an input image including a main subject and a background, and generating a first image including the main subject and a second image including the background from the input image; a generating step of generating text data indicating characteristics of the background as a prompt based on the second image; An image processing method comprising:

26. A program for causing a computer to function as each of the means of the image processing apparatus according to claim 1.

Citation Information

Patent Citations

  • Solid state image sensor, imaging apparatus, and information processing system

    JP2022018997A