Image processing device, image processing method, program, and imaging system
The image processing device and method address the challenge of generating high-quality images with latent space analysis by performing multi-layer hierarchical encoding and decoding, facilitating efficient analytical information generation in embedded systems.
Patent Information
- Application Number
- PCT/JP2025/023396
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-16
- Filing Date
- 2025-06-30
- Publication Date
- 2026-01-22
AI Technical Summary
Existing image generation models using diffusion technology struggle to generate high-quality images without altering their content by inputting analytical information representing multiple types of analysis results in a latent space, and there is no method for easily generating multiple types of analytical information for images in a latent space.
An image processing device and method that converts an image in pixel space into latent space, performing multi-layer hierarchical encoding and decoding to generate decoded results as analysis information representing different types of analysis in the latent space, using a combination of encoding and decoding results from adjacent layers.
Enables easy generation of multiple types of analytical information in latent space with reduced computational cost, suitable for implementation in embedded systems like imaging devices.
Smart Images

Figure JP2025023396_22012026_PF_FP_ABST
Abstract
Description
Image processing device, image processing method, program, and imaging system
[0001] The present technology relates to an image processing device, an image processing method, a program, and an imaging system, and in particular to an image processing device, an image processing method, a program, and an imaging system that enable easy generation of multiple types of analytical information of an image in a latent space.
[0002] Recently, diffusion technology has become popular as an image generation AI (artificial intelligence) technology. This technology generates images by machine learning Gaussian dediffusion, a denoising process, from an image to which Gaussian noise has been added, and then repeatedly executing this Gaussian dediffusion process. This diffusion technology has significantly higher image generation capabilities than other generative AI technologies, such as VAE (variational autoencoder) and GAN (generative adversarial network). Therefore, methods using this diffusion technology to generate image generation models, such as foundation models for multimodal image generation AI that integrate and process text and music, are being developed around the world.
[0003] In such an image generation model, for example, a user inputs text (characters) as a prompt specifying the detailed composition of an image, such as "A white horse is running in a meadow. Three brown horses are in the background. There is a forest and a brick wall in the background." In this case, an image corresponding to this composition is generated, but since it is difficult for the user to specify the exact composition, it is difficult to generate an image with the user's desired composition. Note that the user can specify a composition that cannot be specified by the prompt using a function such as a Stable Diffusion control net, but even using this function, it is difficult to generate an image with the user's desired composition.
[0004] A user can also generate moving images by inputting low-resolution (e.g., 32 x 24 pixel) moving images, such as proxy images, as prompts. In this case, the user does not need to input prompts for each image, and can specify the exact composition in the prompts. However, because the moving images input as prompts have low resolution and little information, the image generation capability is lower than when text prompts are input specifying an object, such as a cat or a landscape.
[0005] Image-to-image processing can also be performed using an image generation model that uses diffusion technology to generate an output image that is modified based on an input image. However, in this case, detailed parts of the image may be modified, and an image may be generated that includes, for example, a false subject (object) that does not exist in the original image. Therefore, such an image generation model is not suitable for cases where a content creator does not want to change the content of the base image, such as when generating a captured image that has undergone predetermined camera signal processing, such as controlling the denoising level, from a captured image.
[0006] Image generation models using diffusion technology use U-Net, which is effective for image segmentation analysis, depth analysis, etc. U-Net is used together with VAE and VQ-VAE (Vector Quantization Variational Autoencoder) as a method capable of high-speed processing of high-frequency data to low-frequency data. For example, super-resolution technology using U-Net has been devised (see, for example, Patent Document 1).
[0007] US Patent Application Publication No. 2023 / 0267652
[0008] However, in image generation models, it has not been considered to improve image generation capabilities and generate high-quality images without altering their content by inputting analytical information representing multiple types of analysis results of images in a latent space as prompts. Therefore, no method has been considered for easily generating multiple types of analytical information for images in a latent space. Note that a latent space is a feature space that simplifies data representation.
[0009] The present technology has been made in consideration of such circumstances, and makes it possible to easily generate multiple types of analytical information for images in latent space.
[0010] An image processing device according to a first aspect of the present technology includes an image encoding unit that converts an image in pixel space into an image in latent space, and an image analysis unit that performs multi-layer hierarchical encoding on the image in the latent space and performs multi-layer hierarchical decoding on the encoding result of the lowest layer of the multiple layers, thereby generating decoded results of each of the multiple layers as analysis information representing the results of different types of analysis of the image in the latent space, and the image analysis unit is configured to decode each of the multiple layers using information that combines the encoding result of that layer with the decoding result of the layer one layer below it.
[0011] An image processing method according to a first aspect of the present technology includes converting an image in pixel space into an image in latent space, performing multi-layer hierarchical encoding on the image in the latent space, and performing multi-layer hierarchical decoding on the encoding result of the lowest layer of the multiple layers, thereby generating decoded results of each of the multiple layers as analysis information representing the results of different types of analysis of the image in the latent space, wherein the decoding of each of the multiple layers is performed on information combining the encoding result of that layer and the decoding result of the layer one layer below that layer.
[0012] A program according to a first aspect of the present technology includes converting an image in pixel space into an image in latent space, performing multi-layer hierarchical encoding on the image in the latent space, and performing multi-layer hierarchical decoding on the encoding result of the lowest layer of the multiple layers, thereby generating decoded results of each of the multiple layers as analysis information representing the results of different types of analysis of the image in the latent space, wherein the decoding of each of the multiple layers is a program for causing a computer to execute processing on information combining the encoding result of that layer and the decoding result of the layer one layer below that layer.
[0013] In a first aspect of the present technology, an image in pixel space is converted into an image in latent space, the image in the latent space is subjected to multi-layer hierarchical encoding, and the encoding result of the lowest layer of the multi-layer is subjected to multi-layer hierarchical decoding, whereby the decoding result of each of the multi-layer is generated as analysis information representing the result of a different type of analysis of the image in the latent space. Note that the decoding of each of the multi-layer is performed on information combining the encoding result of that layer and the decoding result of the layer immediately below it.
[0014] An imaging system according to a second aspect of the present technology includes an imaging unit that captures an image in pixel space, an image encoding unit that converts the image in pixel space captured by the imaging unit into an image in latent space, and an image analysis unit that performs multi-layer hierarchical encoding on the image in latent space and performs multi-layer hierarchical decoding on the encoding result of the lowest layer of the multiple layers, thereby generating decoded results of each of the multiple layers as analysis information representing the results of different types of analysis of the image in the latent space, wherein the image analysis unit is configured to decode each of the multiple layers using information that combines the encoding result of that layer with the decoding result of the layer one layer below it.
[0015] In a second aspect of the present technology, an image in pixel space is captured, the captured image in pixel space is converted into an image in latent space, a multi-layer hierarchical encoding is performed on the image in latent space, and the multi-layer hierarchical decoding is performed on the encoding result of the lowest layer of the multiple layers, thereby generating a decoding result of each of the multiple layers as analysis information representing a result of a different type of analysis of the image in the latent space. Note that the decoding of each of the multiple layers is performed on information combining the encoding result of that layer and the decoding result of the layer one layer below that layer.
[0016] 16 is a diagram illustrating an example of an overview of image analysis using CNN. FIG. 17 is a block diagram illustrating an example of the configuration of a first embodiment of an imaging system to which the present technology is applied. FIG. 18 is a block diagram illustrating an example of the configuration of the image analysis unit of FIG. 2. FIG. 19 is a diagram illustrating an overview of the processing of the image analysis unit of FIG. 3. FIG. 20 is a diagram illustrating examples of expected values of an input image and analysis meta information. FIG. 21 is a diagram illustrating examples of inferred values of a restored image and analysis meta information. FIG. 22 is a diagram illustrating expected values of an input image, a restored image, and analysis meta information, and examples of inferred values of an input image with reduced resolution, a restored image, and analysis meta information. FIG. 23 is a flowchart illustrating image analysis processing. FIG. 24 is a block diagram illustrating an example of the configuration of a learning device that machine-learns an image analysis model. FIG. 25 is a flowchart illustrating learning processing by the learning device of FIG. 9. FIG. 26 is a diagram illustrating an overview of an example of a learning method for an image analysis model. FIG. 27 is a block diagram illustrating an example of the configuration of a second embodiment of an imaging system to which the present technology is applied. FIG. 28 is a block diagram illustrating an example of the configuration of the image processing unit of FIG. 22. FIG. 29 is a block diagram illustrating an example of the configuration of the image generation unit of FIG. 23. FIG. 29 is a flowchart illustrating crop interpolation processing. FIG. 29 is a block diagram illustrating an example of the configuration of a learning device that machine-learns a crop interpolation model. FIG. 29 is a flowchart illustrating learning processing by the learning device of FIG. 16. FIG. 29 is a diagram illustrating an overview of an example of a learning method for a crop interpolation model. 25 is a diagram showing an example of a learning method for learning an IMU analysis model; FIG. 26 is a diagram showing an example of a learning device for machine learning an IMU analysis model; FIG. 27 is a diagram showing an example of a learning method for learning an IMU analysis model; FIG. 28 is a diagram showing an example of a learning device for machine learning an IMU analysis model; FIG. 29 is a diagram showing an example of a learning method for learning an IMU analysis model; FIG. 30 is a diagram showing an example of a learning method for learning an IMU analysis model; FIG. 31 is a diagram showing an example of a learning method for learning an IMU analysis model; FIG. 32 is a diagram showing an example of a learning method for learning an IMU analysis model; FIG. 33 is a diagram showing an example of a learning method for learning an IMU analysis model; FIG. 34 is a diagram showing an example of a learning method for learning an IMU analysis model; FIG. 35 is a diagram showing an example of a learning method for learning an IMU analysis model; FIG. 36 is a diagram showing an example of a learning method for learning an IMU analysis model;Fig. 31 is a block diagram showing a configuration example of a sixth embodiment of an imaging system to which the present technology is applied. Fig. 32 is a block diagram showing a configuration example of the image generation unit of Fig. 30. Fig. 33 is a diagram explaining the effect of light crop interpolation processing. Fig. 34 is a diagram showing an example of segment analysis information. Fig. 35 is a diagram showing an example of a bintree structure of a segment. Fig. 36 is a block diagram showing a configuration example of computer hardware.
[0017] Hereinafter, modes for carrying out the present technology (hereinafter referred to as embodiments) will be described. The description will be made in the following order: 0. Example of an overview of image analysis using CNN 1. First embodiment (imaging system that performs image analysis processing) 2. Second embodiment (imaging system that performs crop interpolation processing) 3. Third embodiment (imaging system that performs IMU analysis processing) 4. Fourth embodiment (imaging system that performs occlusion interpolation processing) 5. Fifth embodiment (image processing system) 6. Sixth embodiment (imaging system that performs crop interpolation processing with a light load) 7. Extended function of image analysis model 8. Computer
[0018] <0. Example of Overview of Image Analysis Using CNN> FIG. 1 is a diagram for explaining an example of an overview of segment analysis, depth analysis, and motion analysis using a CNN (Convolutional Neural Network).
[0019] When segment analysis, depth analysis, and motion analysis are performed using CNN, a dedicated network is created for each analysis task, and each analysis is performed using each network.
[0020] Specifically, as shown in Figure 1, a network dedicated to segmentation analysis is used to repeatedly perform convolution (Conv) and pooling on an image layer by layer, and the calculation results of each layer are propagated to the next layer using an appropriate firing function. This allows the structure within the image to be understood in the order of edges, parts, objects, and faces, for example. The output of the most abstract layer is then connected to a fully connected layer and labeled to output analytical information from the segmentation analysis.
[0021] As with the analysis information of the segment analysis, the analysis information of the depth analysis is output using a network dedicated to the depth analysis, and the analysis information of the motion analysis is output using a network dedicated to the motion analysis.
[0022] Although the analytical information generated in this way is highly accurate, it is computationally expensive and requires a great deal of resources to generate the analytical information. Therefore, it is difficult to perform image analysis using CNN on a digital signal processor (DSP) in an embedded system.
[0023] 1. First Embodiment Configuration Example of Imaging System FIG. 2 is a block diagram showing a configuration example of a first embodiment of an imaging system to which the present technology is applied.
[0024] 2 includes an imaging unit 11, an input unit 12, an image processing unit 13, and a storage unit 14. The imaging system 10 generates, from a moving image made up of RGB images of multiple frames in pixel space captured by the imaging unit 11, multiple types of analytical information and the like of the RGB images of each frame in pixel space.
[0025] Specifically, the imaging unit 11 captures images and supplies the resulting pixel space moving images for each frame as pixel space input images to the image processing unit 13 and the storage unit 14. The input unit 12 acquires text data input as a prompt from the user, which is used to specify the optical depth, the degree of motion blur, etc., and supplies the text data to the image processing unit 13.
[0026] The image processing unit 13 is composed of a reduction unit 21, a VAE encoder 22, a text encoder 23, an image analysis unit 24, a VAE decoder 25, and decoders 26 to 28. For each frame, the image processing unit 13 uses an image analysis model, which is a learning model of the image analysis AI, to generate multiple types of analysis information in the pixel space of the input image from the input image in pixel space, and performs image analysis processing to remove optical noise. Optical noise includes crop areas, flicker, motion blur, etc.
[0027] The reduction unit 21 performs down-conversion processing, pixel thinning processing, etc. on the input image (high-resolution image) in pixel space supplied from the imaging unit 11, thereby reducing the resolution of the input image in pixel space. For example, the reduction unit 21 reduces the resolution of the input image in pixel space of 1920 x 1080 pixels supplied from the imaging unit 11 to an input image in pixel space of 32 x 24 pixels. The reduction unit 21 supplies the input image in pixel space with reduced resolution to the VAE encoder 22.
[0028] The VAE encoder 22 (image encoding unit) converts the input image in pixel space supplied from the reduction unit 21 into an image in latent space by performing convolution processing and downsizing processing on the input image in pixel space. The VAE encoder 22 supplies the image in latent space to the image analysis unit 24.
[0029] The text encoder 23 converts the text data in the character space supplied from the input unit 12 into text data in the latent space and supplies it to the image analysis unit 24 .
[0030] The image analysis unit 24 uses an image analysis model having a U-Net structure to generate multiple types of analysis information of the latent space of an image from the latent space image supplied from the VAE encoder 22. Specifically, the image analysis unit 24 performs three-layer hierarchical encoding on the latent space image and three-layer hierarchical decoding on the encoding result of the lowest layer of the three layers. As a result, the image analysis unit 24 generates the decoding results of each of the three layers as analysis information of latent space segment analysis, analysis information of depth analysis, and analysis information of motion analysis, respectively. Note that when performing hierarchical encoding and hierarchical decoding, the image analysis unit 24 also uses text data of the latent space input from the text encoder 23 as necessary.
[0031] The image analysis unit 24 combines segment analysis information, which is analysis information of the segment analysis of the generated latent space, with the image of the latent space from the VAE encoder 22, to generate a restored image in the latent space by removing optical noise from the image. The image analysis unit 24 supplies the restored image in the latent space to the VAE decoder 25. The image analysis unit 24 supplies the generated segment analysis information of the latent space to the decoder 26. The image analysis unit 24 supplies depth analysis information, which is analysis information of the depth analysis of the generated latent space, to the decoder 27, and motion analysis information, which is analysis information of the motion analysis of the latent space, to the decoder 28.
[0032] The VAE decoder 25 (image decoding unit) decodes the restored image in the latent space supplied from the image analysis unit 24, and generates a restored image in the pixel space with the same resolution as the resolution of the input image whose resolution has been reduced by the reduction unit 21. The VAE decoder 25 supplies the restored image in the pixel space to the storage unit 14.
[0033] The decoder 26 (analysis information decoding unit) converts the segment analysis information in the latent space supplied from the image analysis unit 24 into segment analysis information in the pixel space that represents the ground, the sky, or something other than the ground and the sky as a segment of each pixel of the low-resolution input image. The decoder 26 supplies the segment analysis information in the pixel space to the storage unit 14.
[0034] The decoder 27 (analysis information decoding unit) converts the depth analysis information in the latent space supplied from the image analysis unit 24 into depth analysis information in the pixel space that represents the depth of each pixel of the input image whose resolution has been reduced. The decoder 27 supplies the depth analysis information in the pixel space to the storage unit 14.
[0035] The decoder 28 (analysis information decoding unit) converts the motion analysis information in the latent space supplied from the image analysis unit 24 into motion analysis information in the pixel space, which represents the optical flow of each pixel of the input image with reduced resolution. The decoder 28 supplies the motion analysis information in the pixel space to the storage unit 14.
[0036] As described above, the image processing unit 13 generates segment analysis information, depth analysis information, and motion analysis information by capturing the entire composition of the input image, the resolution of which has been reduced by the reduction unit 21, from the perspective of histogram analysis using the image analysis model. The image processing unit 13 reduces the resolution of the input image in pixel space to generate a restored image, segment analysis information, depth analysis information, and motion analysis information, and therefore can reduce calculation costs compared to when the resolution of the input image in pixel space is not reduced.
[0037] The storage unit 14 stores the restored image in pixel space of each frame, the segment analysis information, the depth analysis information, and the motion analysis information as analysis meta information in association with the input image in pixel space of each frame supplied from the imaging unit 11. This analysis meta information is used to control the imaging unit 11 to realize, for example, an autofocus function, a panning function for tracking and photographing a moving subject, and the like.
[0038] The components constituting the imaging system 10 may be provided in the same device or in separate devices. For example, the imaging system 10 may be an imaging device including an imaging unit 11, an input unit 12, an image processing unit 13, and a storage unit 14. The imaging system 10 may also be an imaging system including an imaging device including the imaging unit 11 and an image processing device including the input unit 12, the image processing unit 13, and the storage unit 14. The image processing unit 13 does not need to be provided with the reduction unit 21.
[0039] <Configuration Example of Image Analysis Unit> FIG. 3 is a block diagram showing a configuration example of the image analysis unit 24 in FIG.
[0040] 3 has a U-Net structure. Specifically, the image analysis unit 24 includes a segment encoder 41, a depth encoder 42, and a motion encoder 43. The image analysis unit 24 also includes a motion decoder 44, a combining unit 45, a depth decoder 46, a combining unit 47, a segment decoder 48, and a combining unit 49.
[0041] The segment encoder 41 is an encoder that encodes the segment layer, which is the highest layer of the three layers, and has a residual connection block including a convolutional layer and a self-attention mechanism block. The segment encoder 41 is skip-connected to the segment decoder 48 via the connection unit 47. The segment encoder 41 encodes the latent space image provided by the VAE encoder 22 and provides the encoded result to the depth encoder 42 and the connection unit 47.
[0042] The depth encoder 42 is an encoder that encodes the depth layer, which is one layer below the segment layer, and includes a residual combination block and a self-attention mechanism block. The depth encoder 42 is skip-connected to the depth decoder 46 via a combination unit 45. The depth encoder 42 encodes the encoding result provided by the segment encoder 41 and provides the encoding result to the motion encoder 43 and the combination unit 45.
[0043] The motion encoder 43 is an encoder that encodes the motion layer, which is the lowest of the three layers, and includes a residual combination block and a self-attention mechanism block. The motion encoder 43 is skip-connected to the motion decoder 44. The motion encoder 43 encodes the encoding results of three consecutive frames, which are supplied from the depth encoder 42 and consist of the frame to be processed and the frames before and after that frame.
[0044] In this way, the motion encoder 43 encodes each frame using the encoding results of three consecutive frames including that frame. Therefore, in order to reduce the computational cost of the image analysis process, the VAE encoder 22, segment encoder 41, and depth encoder 42 perform processing using pipeline processing. Note that the number of frames corresponding to the encoding results used to encode each frame by the motion encoder 43 is not limited to three. The motion encoder 43 supplies the encoding results to the motion decoder 44.
[0045] The motion decoder 44 is a decoder that performs decoding of the motion layer and includes a residual combination block and a self-attention mechanism block. The motion decoder 44 decodes the encoding result provided by the motion encoder 43. The motion decoder 44 provides the decoding result to the combination unit 45 and also provides it to the decoder 28 in Figure 2 as motion analysis information in the latent space.
[0046] The combining unit 45 combines the encoding result supplied from the depth encoder 42 with the decoding result supplied from the motion decoder 44 , and supplies the resulting information to a depth decoder 46 .
[0047] The depth decoder 46 is a decoder that performs decoding of the depth layer and includes a residual combination block and a self-attention mechanism block. The depth decoder 46 decodes the information provided by the combination unit 45. The depth decoder 46 provides the decoded result to the combination unit 47 and also provides it to the decoder 27 in FIG. 2 as depth analysis information in the latent space.
[0048] The combining unit 47 combines the encoding result supplied from the segment encoder 41 with the decoding result supplied from the depth decoder 46 , and supplies the resulting information to a segment decoder 48 .
[0049] The segment decoder 48 is a decoder that performs decoding of the segment layer and includes a residual combination block and a self-attention mechanism block. The segment decoder 48 decodes the information provided by the combination unit 47. The segment decoder 48 provides the decoded result to the combination unit 49 and also provides it to the decoder 26 in Figure 2 as segment analysis information in the latent space.
[0050] The combining unit 49 combines the image in the latent space supplied from the VAE encoder 22 with the decoding result supplied from the segment decoder 48, and supplies the resulting information to the VAE decoder 25 in Figure 2 as a restored image in the latent space.
[0051] The image analysis unit 24 does not necessarily need to include a self-attention block. The self-attention block improves output accuracy, but requires memory. Therefore, if the image processing unit 13 is implemented as an embedded system, it is preferable not to include a self-attention block. The information combined with the encoding result in the combining unit 45 (47, 49) may not be the decoding result of the lower layer, but may be an intermediate result of the decoding of the lower layer.
[0052] As described above, the image analysis unit 24 incorporates an encoder and decoder using an autoencoder-like network that is lighter than a CNN at each layer. The tasks of generating latent space segment analysis information, depth analysis information, and motion analysis information are then realized at each layer below the top layer. That is, latent space analysis meta-information is generated as an intermediate product at each layer below the top layer. Therefore, the computational cost of image analysis is reduced compared to when image analysis is performed using a CNN as described in Figure 1. This effect is particularly useful when the image analysis unit 24 is implemented using a DSP in an embedded system.
[0053] The depth decoder 46 generates depth analysis information in the latent space using motion analysis information in the latent space, which is the decoded result of the motion decoder 44, thereby improving the accuracy of the depth analysis information in the latent space. The segment decoder 48 generates segment analysis information in the latent space using the depth analysis information in the latent space, thereby improving the accuracy of the segment analysis information in the latent space. The VAE decoder 25 generates a reconstructed image in the pixel space using the segment analysis information in the latent space, thereby improving the inference accuracy of the reconstructed image.
[0054] <Outline of Processing by Image Analysis Unit> FIG. 4 is a diagram for explaining an outline of processing by the image analysis unit 24 in FIG.
[0055] 4, in the image analysis unit 24, the latent space image supplied from the VAE encoder 22 is encoded by the segment encoder 41, and the encoding result is pooled and encoded by the depth encoder 42. The encoding result is pooled and encoded by the motion encoder 43.
[0056] The encoding result by the motion encoder 43 is decoded by the motion decoder 44 and output as motion analysis information in the latent space. This latent space motion analysis information is combined (plugged in) with the encoding result by the depth encoder 42 by the combining unit 45, decoded by the depth decoder 46, and output as depth analysis information in the latent space. This latent space depth analysis information is combined with the encoding result by the segment encoder 41 by the combining unit 47, decoded by the segment decoder 48, and output as segment analysis information in the latent space.
[0057] <Examples of expected values and inferred values of analysis meta information> Fig. 5 is a diagram showing an example of an input image in pixel space and an expected value of analysis meta information for the input image. Fig. 6 is a diagram showing an example of a restored image in pixel space and analysis meta information output from the image processing unit 13 when the input image of Fig. 5 is input to the image processing unit 13.
[0058] 5A to 5D respectively represent the input image in pixel space, the expected value of the segment analysis information, the expected value of the depth analysis information, and the expected value of the motion analysis information. In the example of FIG. 5, the resolution of the input image in pixel space is 1920 x 1080 pixels. Therefore, the expected value of the segment analysis information in pixel space, the expected value of the depth analysis information, and the expected value of the motion analysis information are each composed of information representing the segment, depth, and optical flow of each of the 1920 x 1080 pixels.
[0059] The expected values of the segment analysis information, the expected values of the depth analysis information, and the expected values of the motion analysis information are generated from the input image in the pixel space of A in Fig. 5 using CNN with a GPU (Graphics Processing Unit) or the like. The expected values of the segment analysis information, the expected values of the depth analysis information, and the expected values of the motion analysis information are used as training data when an image analysis model is machine-learned.
[0060] The expected values of the depth analysis information and the expected values of the motion analysis information may be actual measurements acquired by a depth sensor and a motion sensor when capturing the input image of Fig. 5A. A program for generating the expected values of the segment analysis information, the expected values of the depth analysis information, and the expected values of the motion analysis information can be obtained from OpenCV (Open Source Computer Vision Library) or the like. This process requires high computational costs and is difficult to execute in an embedded system.
[0061] 6A to 6D respectively show a restored image in pixel space, segment analysis information (inferred values), depth analysis information (inferred values), and motion analysis information (inferred values). In the example of FIG. 6, the resolution of the restored image is 32 × 24 pixels. Therefore, the segment analysis information, depth analysis information, and motion analysis information in pixel space each consist of information representing the segment, depth, and optical flow of each of the 32 × 24 pixels.
[0062] Although the resolution of the inferred values of the analysis meta information of B to D in Fig. 6 is lower than the resolution of the expected values of the analysis meta information of B to D in Fig. 5, they are generated roughly as expected. Therefore, this segment analysis information, depth analysis information, and motion analysis information can be used for camera signal processing of the input image in pixel space.
[0063] <Other examples of expected values and inferred values of analysis meta information> Figure 7 is a diagram showing examples of an input image in pixel space, the expected values of the restored image and analysis meta information corresponding to the input image, an input image whose resolution has been reduced by the reduction unit 21, and the restored image and analysis meta information output from the image processing unit 13.
[0064] 7A to 7C, the images in the first row of the 2 (rows) x 3 (columns) images represent, from left to right, an input image in pixel space, an input image before optical noise corresponding to the expected value of the restored image is added, and a black image. The images in the second row represent, from left to right, the expected values of segment analysis information, depth analysis information, and motion analysis information of the input image in pixel space. In the examples of FIGS. 7A to 7C, the input images in pixel space are, respectively, an image without optical noise, an image with a cropped region as optical noise, and an image with both a cropped region and flicker as optical noise.
[0065] Each of D to F in Fig. 7 corresponds to each of A to C in Fig. 7. The 2 (rows) x 3 (columns) images in D to F in Fig. 7 represent, from left to right, a reduced-resolution input image when any of the corresponding input images in A to C in Fig. 7 is input to the image processing unit 13, an expected value of the restored image, and (the inferred value of) the restored image output from the image processing unit 13. The images in the second row represent, from left to right, segment analysis information (the inferred value), depth analysis information (the inferred value), and motion analysis information (the inferred value) output from the image processing unit 13. Note that each of the images in D to F in Fig. 7 is surrounded by a black frame.
[0066] As shown in Figure 7, regardless of the presence or absence and type of optical noise in the input image in pixel space, the restored images, segment analysis information, depth analysis information, and motion analysis information shown in Figures 7D to 7F are generated generally as expected.
[0067] <Explanation of Image Analysis Processing> Fig. 8 is a flowchart illustrating the image analysis processing by the image processing unit 13 in Fig. 2. This image analysis processing is started when an input image in pixel space is input from the imaging unit 11, for example.
[0068] 8 , the reduction unit 21 reduces the resolution of the input image in pixel space supplied from the imaging unit 11, and supplies the reduced image to the VAE encoder 22. In step S11, the VAE encoder 22 converts the input image, the resolution of which has been reduced by the processing in step S10, into an image in latent space, and supplies the image to the image analysis unit 24.
[0069] In step S12, the image analysis unit 24 generates latent space analysis meta information for the image from the latent space image obtained by the processing of step S11 using the image analysis model. At this time, if text data is input from the input unit 12 to the image processing unit 13, the latent space text data converted from the text data by the text encoder 23 is used as needed. The image analysis unit 24 supplies information combining the generated latent space segment analysis information and the latent space image to the VAE decoder 25. The image analysis unit 24 supplies the generated latent space segment analysis information to the decoder 26, depth analysis information to the decoder 27, and motion analysis information to the decoder 28.
[0070] In step S13, the VAE decoder 25 decodes the information supplied from the image analysis unit 24 to generate a restored image in pixel space, and supplies the image to the storage unit 14. In step S14, the decoder 26 converts the segment analysis information in the latent space generated by the processing in step S12 into segment analysis information in pixel space, and supplies the converted information to the storage unit 14.
[0071] In step S15, the decoder 27 converts the depth analysis information in the latent space generated by the processing in step S12 into depth analysis information in the pixel space, and supplies the converted information to the storage unit 14. In step S16, the decoder 28 converts the motion analysis information in the latent space generated by the processing in step S12 into motion analysis information in the pixel space, and supplies the converted information to the storage unit 14.
[0072] In step S17, the storage unit 14 stores the restored image in pixel space and the analysis meta information corresponding to the input image in pixel space, in association with the input image in pixel space supplied from the imaging unit 11. Then, the image analysis process ends.
[0073] As described above, in the image processing unit 13, the VAE encoder 22 converts the low-resolution input image in pixel space into an image in latent space. The image analysis unit 24 performs three-layer hierarchical encoding on the latent space image and three-layer hierarchical decoding on the encoding result of the lowest layer, thereby generating the decoded results of each layer as analytical meta information in the latent space. Note that the image analysis unit 24 decodes each layer using information combining the encoding result of that layer with the decoding result of the layer immediately below it. Therefore, the image processing unit 13 can easily generate analytical meta information in the latent space from the low-resolution input image in pixel space with low computational cost. As a result, the image processing unit 13 can be implemented using a DSP or the like in an embedded system such as an imaging device.
[0074] <Configuration Example of Learning Device> FIG. 9 is a block diagram showing a configuration example of a learning device that learns an image analysis model by machine learning.
[0075] The learning device 50 in FIG. 9 is made up of an acquisition unit 60, a reduction unit 61, an optical noise addition unit 62, a learning unit 63, reduction units 64 to 66, and loss calculation units 67 to 70.
[0076] The acquisition unit 60 acquires and stores a training set consisting of training data for multiple frames from a server (not shown). This training data consists of an ideal, high-quality pixel space image for each frame and expected values of analytical meta information for that image. This training data also includes text data in character space corresponding to the image as needed.
[0077] The acquisition unit 60 sequentially reads out, frame by frame, the training data for multiple frames included in the training set it holds. The acquisition unit 60 outputs the image in pixel space and the text data in character space from the read training data to the reduction unit 61 and the training unit 63, respectively. The acquisition unit 60 supplies the expected values of the segment analysis information, the expected values of the depth analysis information, and the expected values of the motion analysis information from the read training data to the reduction units 64 to 66, respectively.
[0078] The reduction unit 61 reduces the resolution of the image in pixel space supplied from the acquisition unit 60 in the same manner as the reduction unit 21 in Fig. 2. For example, the reduction unit 61 reduces the resolution of an image in pixel space of 1920 x 1080 pixels to an image in pixel space of 32 x 24 pixels. The reduction unit 61 supplies the reduced-resolution image to the optical noise addition unit 62 and the loss calculation unit 67.
[0079] The optical noise adding unit 62 adds optical noise that is assumed to have been added to the input image to the image supplied from the reducing unit 61, and supplies the resulting image to the learning unit 63. This cropped area is, for example, a black area that does not contain image information, and is added by electronic stabilization processing (electronic image stabilization processing). This electronic stabilization processing is performed by, for example, the imaging unit 11.
[0080] The learning unit 63 uses an image analysis model to generate a restored image in pixel space and analysis meta information from the image supplied from the optical noise adder 62. Alternatively, the learning unit 63 uses a learning model corresponding to the image analysis model to generate, from the image supplied from the optical noise adder 62, an inferred value of the optical noise added to the image in pixel space and analysis meta information. Note that, hereinafter, the process by which the learning unit 63 generates a restored image in pixel space or an inferred value of the optical noise and analysis meta information from the image supplied from the optical noise adder 62 is referred to as image analysis forward processing. During the image analysis forward processing, the learning unit 63 uses text data supplied from the acquisition unit 60 as necessary. The learning unit 63 supplies the restored image in pixel space or the inferred value of the optical noise, segment analysis information, depth analysis information, and motion analysis information to each of the loss calculation units 67 to 70.
[0081] Time-embedded information (TimeEmbedded), which is connected to the residual combination block and controls the time during denoising, is input to the learning unit 63. This time-embedded information is used for control based on time information in repeated execution of noise dediffusion when the learning unit 63 performs error propagation using a denoising loss value.
[0082] The reduction unit 64 reduces the expected value of the segment analysis information supplied from the acquisition unit 60 to segment analysis information of the same resolution as the image whose resolution has been reduced by the reduction unit 61, and supplies the reduced value to a loss calculation unit 68. The reduction unit 65 reduces the expected value of the depth analysis information supplied from the acquisition unit 60 to depth analysis information of the same resolution as the image whose resolution has been reduced by the reduction unit 61, and supplies the reduced value to a loss calculation unit 69. The reduction unit 66 reduces the expected value of the motion analysis information supplied from the acquisition unit 60 to motion analysis information of the same resolution as the image whose resolution has been reduced by the reduction unit 61, and supplies the reduced value to a loss calculation unit 70.
[0083] The loss calculation unit 67 calculates a reconstruction loss value using the image supplied from the reduction unit 61 as an expected value and the restored image supplied from the learning unit 63 as an inferred value. Alternatively, the loss calculation unit 67 generates, as an expected value of optical noise, the difference between the image from the reduction unit 61 to which no optical noise has been added and the image from the optical noise addition unit 62 to which optical noise has been added. The loss calculation unit 67 then calculates a denoising loss value using the expected value of optical noise and the inferred value of optical noise from the learning unit 63. The loss calculation unit 67 performs machine learning of the image analysis model by propagating an error to the learning unit 63 using the calculated reconstruction loss value or denoising loss value. This updates the parameters (weights) of the image analysis model.
[0084] When the reconstruction loss value is error-propagated, the learning unit 63 can machine-learn an image analysis model that generates a restored image that is faithful to the input image. Therefore, the image analysis model machine-learned in this manner is suitable for camera signal processing, etc. On the other hand, when the denoising loss value is error-propagated, an image analysis model that can generate a restored image that includes information that is not present in the input image can be machine-learned. The calculation cost of the denoising loss value is higher than the calculation cost of the reconstruction loss value.
[0085] The loss calculation unit 68 calculates a reconstruction loss value using the expected value of the segment analysis information supplied from the reduction unit 64 and (the inferred value of) the segment analysis information supplied from the learning unit 63. The loss calculation unit 68 performs machine learning of the image analysis model by propagating an error to the learning unit 63 using the reconstruction loss value. As a result, the parameters of the image analysis model are updated.
[0086] The loss calculation unit 69 calculates a reconstruction loss value using the expected value of the depth analysis information supplied from the reduction unit 65 and (the inferred value of) the depth analysis information supplied from the learning unit 63. The loss calculation unit 69 performs machine learning of the image analysis model by propagating an error to the learning unit 63 using the reconstruction loss value. As a result, the parameters of the image analysis model are updated.
[0087] The loss calculation unit 70 calculates a reconstruction loss value using the expected value of the motion analysis information supplied from the reduction unit 66 and (the inferred value of) the motion analysis information supplied from the learning unit 63. The loss calculation unit 70 performs machine learning of the image analysis model by error-propagating the reconstruction loss value to the learning unit 63. As a result, the parameters of the image analysis model are updated.
[0088] The loss calculation units 67 to 70 weight the reconstruction loss value or the denoising loss value as needed.
[0089] <Explanation of Learning Process> FIG. 10 is a flowchart illustrating the learning process performed by the learning device 50 of FIG.
[0090] 10, the learning device 50 initializes the hyperparameters of the image analysis model. In step S32, the learning device 50 initializes the network definition of the image analysis model. In step S33, the acquisition unit 60 initializes the currently held learning set.
[0091] In step S34, the acquisition unit 60 reads (loads) the learning data of frames that have not yet been read out from the learning data of each of the multiple frames included in the learning set initialized by the processing of step S33. The acquisition unit 60 supplies the image in pixel space and the text data in character space of the learning data to the reduction unit 61 and the learning unit 63, respectively. The acquisition unit 60 supplies the expected values of the segment analysis information, the expected values of the depth analysis information, and the expected values of the motion analysis information of the read learning data to the reduction units 64 to 66, respectively.
[0092] In step S35, the reduction unit 61 reduces the resolution of the image in pixel space supplied from the acquisition unit 60, and supplies the reduced resolution image to the optical noise addition unit 62 and the loss calculation unit 67. In step S36, the optical noise addition unit 62 adds optical noise to the image whose resolution has been reduced by the processing in step S35, and supplies the reduced resolution image to the learning unit 63.
[0093] In step S37, the learning unit 63 performs image analysis forward processing and supplies the resulting restored image in pixel space or the estimated value of optical noise, segment analysis information, depth analysis information, and motion analysis information to the loss calculation units 67 to 70, respectively.
[0094] In step S38, the reduction unit 64 reduces the expected value of the segment analysis information supplied from the acquisition unit 60 and supplies the resulting segment analysis information, having the same resolution as the image whose resolution has been reduced by the reduction unit 61, to the loss calculation unit 68. The reduction unit 65 reduces the expected value of the depth analysis information supplied from the acquisition unit 60 and supplies the resulting depth analysis information, having the same resolution as the image whose resolution has been reduced by the reduction unit 61, to the loss calculation unit 69. The reduction unit 66 reduces the expected value of the motion analysis information supplied from the acquisition unit 60 and supplies the resulting motion analysis information, having the same resolution as the image whose resolution has been reduced by the reduction unit 61, to the loss calculation unit 70.
[0095] In step S39, the loss calculation unit 67 calculates a reconstruction loss value using the image whose resolution has been reduced by the processing in step S35 and the restored image obtained as a result of the processing in step S37. Alternatively, the loss calculation unit 67 calculates a denoising loss value using the expected value of optical noise and the inferred value of optical noise obtained as a result of the processing in step S37. Note that the expected value of optical noise is the difference between the image whose resolution has been reduced by the processing in step S35 and the image to which optical noise has been added by the processing in step S36. The loss calculation units 68 to 70 each calculate a reconstruction loss value using the expected value reduced by the processing in step S38 and the inferred value obtained as a result of the processing in step S37.
[0096] In step S40, the loss calculation units 67 to 70 use the reconstruction loss value or the denoising loss value calculated in the process of step S39 to propagate errors to the learning unit 63, thereby performing machine learning on the image analysis model. As a result, the parameters of the image analysis model are updated.
[0097] In step S41, the learning device 50 determines whether to end the learning process. Specifically, the learning device 50 determines whether the parameters of the image analysis model have been updated a predetermined number of times (number of epochs) or for a convergence time. If it is determined in step S41 that the learning process should not be ended, that is, if it is determined that the parameters of the image analysis model have not yet been updated a predetermined number of times or for a convergence time, the process returns to step S34, and the subsequent processes are repeated.
[0098] On the other hand, if it is determined in step S41 that the learning process is to be completed, that is, that the parameters of the image analysis model have been updated a predetermined number of times or for a predetermined convergence time, the process proceeds to step S42.
[0099] In step S42, the learning unit 63 outputs (dumps) the parameters of the image analysis model updated in the process of the final step S40 as dictionary data. The image analysis model with this dictionary data set is used in the image analysis unit 24 in Figure 2. After the process of step S42, the learning process ends.
[0100] <Outline of Example of Image Analysis Model Learning Method> FIG. 11 is a diagram illustrating an example of an image analysis model learning method.
[0101] As shown in A of Figure 11, the first learning method is a method of machine learning a foundation model of an image analysis AI as an image analysis model by performing a learning process using a huge learning set 81. In the first learning method, the learning process is performed using the huge learning set 81, so average machine learning is performed for all scenes. The image analysis unit 24 uses the foundation model of the image analysis AI generated by the first learning method as the image analysis model to generate analysis meta information 82 and the like from the video image that is the target of the image analysis process.
[0102] As shown in B of FIG. 11 , the second training method is a method of further fine-tuning the foundation model pre-trained by the first training method using a training set 91 of moving images (video clips) to be subjected to image analysis processing. Fine-tuning is machine learning on a smaller scale than the machine learning by the first training method. The second training method enables machine learning specialized for scenes in the moving images to be subjected to image analysis processing. The image analysis unit 24 generates analysis meta information 92 and the like from the moving images to be subjected to image analysis processing using the fine-tuned image analysis model generated by the second training method.
[0103] As described above, the second learning method performs machine learning specialized for the video scenes to be analyzed. Therefore, the fine-tuned image analysis model is more suitable for the video scenes to be analyzed than the foundation model of the image analysis AI. Therefore, the accuracy of the analysis meta information 92 is higher than that of the analysis meta information 82.
[0104] As for fine-tuning methods, in addition to updating dictionary data through machine learning, there is also a method of efficient additional learning that uses a low-rank matrix called LoRA (Low-Rank Adaptation).
[0105] The types of analysis by the image analysis unit 24 are not limited to segment analysis, depth analysis, and motion analysis. For example, the types of analysis may be ROI (Region of Interest) analysis, flicker noise analysis, etc. A layer that generates analysis information for ROI region analysis and flicker noise analysis in the latent space is provided below the segment layer or depth layer, for example. The number of types of analysis is not limited to three.
[0106] 2. Second Embodiment Configuration Example of Imaging System FIG. 12 is a block diagram showing a configuration example of a second embodiment of an imaging system to which the present technology is applied.
[0107] In the imaging system 110 of Fig. 12, the same reference numerals are used to designate parts corresponding to those in the imaging system 10 of Fig. 2. Therefore, the description of those parts will be omitted as appropriate, and the following description will focus on parts that are different from the imaging system 10.
[0108] 12 includes an input unit 12, an imaging unit 111, an IMU (Inertial Measurement Unit) sensor 112 (inertial sensor), an electronic stabilization processing unit 113, and a VQ-VAE encoder 114. The imaging system 110 also includes an image processing unit 115, a text encoder 116, an image generation unit 117, and a VQ-VAE decoder 118. The imaging system 110 performs electronic stabilization processing in the rotational direction as camera signal processing on moving images captured by the imaging unit 111, and generates a cropless image in which a cropped area generated by the electronic stabilization processing is interpolated.
[0109] The imaging unit 111 captures images and supplies the resulting pixel space moving image (high-resolution image) to the electronic stabilization processing unit 113 for each frame of the image. The IMU sensor 112 (detection unit) is composed of an acceleration sensor (gravity sensor) and an angular velocity sensor (gyro sensor), and detects the acceleration and angular velocity of the imaging unit 111. The IMU sensor 112 supplies IMU information (detection information) representing the acceleration and angular velocity to the electronic stabilization processing unit 113.
[0110] The electronic stabilization processing unit 113 performs electronic stabilization processing in the rotational direction on the image supplied from the imaging unit 11 based on the IMU information supplied from the IMU sensor 112. Camera shake in the rotational direction is corrected in the image after electronic stabilization processing, but a crop region occurs. The electronic stabilization processing unit 113 supplies the image after electronic stabilization processing to the VQ-VAE encoder 114 and also supplies it to the image processing unit 115 as an input image.
[0111] The VQ-VAE encoder 114 converts the pixel space image supplied from the electronic stabilization processor 113 into a latent space image by performing convolution processing and downsizing processing on the pixel space image. The VQ-VAE encoder 114 vector quantizes the latent space image and supplies it to the image generator 117.
[0112] The image processing unit 115 receives text data from the input unit 12. The image processing unit 13 performs latent space image analysis processing, consisting of steps S10 to S12 of the image analysis processing in Fig. 8, on the input image supplied from the electronic stabilization processing unit 113. The image processing unit 115 supplies the latent space input image and restored image as intermediate or final results of the latent space image analysis processing, as well as analysis meta information, to the image generation unit 117.
[0113] The text encoder 116 receives text data in the character space, which is an instruction for generating a cropless image, such as the optical sense of depth of the cropless image, the degree of motion blur, and white balance adjustment based on a more natural sun light source estimation, from the input unit 12. The text encoder 116 converts this text data in the character space into text data in the latent space and supplies it to the image generation unit 117.
[0114] The image generation unit 117 cross-attentionally connects the input image and restored image in the latent space and the analysis meta information supplied from the image processing unit 115 to a crop interpolation model (image generation model), which is a learning model for the AI that generates cropless images. The image generation unit 117 also cross-attentionally connects the text data in the latent space supplied from the text encoder 116 to the crop interpolation model as needed.
[0115] The image generation unit 117 uses this crop interpolation model to perform crop interpolation on the latent space image from the VQ-VAE encoder 114. The crop interpolation process is a process of improving the quality of the latent space image by interpolating the crop region, and generating a cropless image as a high-quality image in the latent space. Note that the resolution of the images before and after the quality improvement is the same. The image generation unit 117 supplies the cropless image in the latent space obtained as a result of the crop interpolation process to the VQ-VAE decoder 118.
[0116] The VQ-VAE decoder 118 decodes the cropless image in latent space supplied from the image generation unit 117 and converts it into a cropless image in pixel space having the same resolution as the video captured by the imaging unit 111. For example, if the resolution of the video captured by the imaging unit 111 is 1920 × 1080 pixels, the resolution of the cropless image in pixel space is also 1920 × 1080 pixels. If the resolution of the video captured by the imaging unit 111 is 1280 × 960 pixels, the resolution of the cropless image in pixel space is also 1280 × 960 pixels. The VQ-VAE decoder 118 outputs the cropless image in pixel space obtained as a result of the conversion.
[0117] It should be noted that instead of the VQ-VAE encoder 114 and the VQ-VAE decoder 118, a VAE encoder and a VAE decoder may be provided.
[0118] <Configuration Example of Image Processing Unit> FIG. 13 is a block diagram showing a configuration example of the image processing unit 115 in FIG.
[0119] In the image processing unit 115 in Figure 13, parts corresponding to the image processing unit 13 in Figure 2 are assigned the same reference numerals. Therefore, explanations of those parts will be omitted as appropriate, and the following explanation will focus on parts that differ from the image processing unit 13. The image processing unit 115 differs from the image processing unit 13 in that it does not include the VAE decoder 25 and decoders 26 to 28, and in that it also outputs the input of layered encoding and the encoding results of each layer. Other than that, it is configured in the same way as the image processing unit 13.
[0120] Specifically, the input image of the latent space input to the image analysis unit 24 is output to the image generation unit 117 in Fig. 12 as the input image of the latent space among the intermediate results of the latent space image analysis process. The encoding results output from the segment encoder 41, the depth encoder 42, and the motion encoder 43 are output to the image generation unit 117 as analysis meta information of the latent space among the intermediate results of the latent space image analysis process.
[0121] The information output from the combining unit 49 is output to the image generating unit 117 as a restored image of the latent space among the final results of the latent space image analysis processing. The decoded results output from the motion decoder 44, the depth decoder 46, and the segment decoder 48 are output to the image generating unit 117 as analysis meta information of the latent space among the final results of the latent space image analysis processing.
[0122] <Configuration Example of Image Generation Unit> FIG. 14 is a block diagram showing a configuration example of the image generation unit 117 in FIG.
[0123] 14 has a U-Net structure. Specifically, the image generation unit 117 is made up of a cross-attention I / F 130, processing units 131 to 137, and combining units 138 to 140.
[0124] The cross-attention I / F 130 supplies the input image and restored image in the latent space and the analysis meta-information supplied from the image processing unit 115 in FIG. 12 to the processing units 131 to 137.
[0125] The processing units 131 to 137 are configured with a residual connection block (ResBlock) and an attention mechanism block (AttnBlock) consisting of a self-attention mechanism and a cross-attention mechanism, and are connected hierarchically. Temporal embedding information is input to the residual connection blocks of the processing units 131 to 137. The input image and restored image in the latent space and analytical meta information supplied from the cross-attention I / F 130 are cross-attention connected to the attention mechanism blocks of the processing units 131 to 137.
[0126] Processing unit 131 is skip-connected to processing unit 137 via coupling unit 140, and processing unit 132 is skip-connected to processing unit 136 via coupling unit 139. Processing unit 133 is skip-connected to processing unit 135 via coupling unit 138.
[0127] The processing units 131 to 133 and part of the processing unit 134 perform three-layer hierarchical encoding, and part of the processing unit 134 and the processing units 135 to 137 perform three-layer hierarchical decoding on the encoding result of the lowest layer.
[0128] Specifically, the processing unit 131 encodes the top layer of the latent space image supplied from the VQ-VAE encoder 114 and supplies the encoded result to the processing unit 132 and the combining unit 140. The processing unit 132 encodes the second-highest layer of the encoded result supplied from the processing unit 131 and supplies it to the processing unit 133 and the combining unit 139. The processing unit 133 encodes the bottom layer of the encoded result supplied from the processing unit 132 and supplies it to the processing unit 134 and the combining unit 138. The processing unit 134 encodes the encoded result supplied from the processing unit 133, decodes the encoded result, and supplies it to the combining unit 138.
[0129] The processing unit 135 decodes the lowest layer of the information supplied from the combining unit 138, which is obtained by combining the encoding result from the processing unit 133 and the decoding result from the processing unit 134, and supplies the decoded result to the combining unit 139. The processing unit 136 decodes the second-highest layer of the information supplied from the combining unit 139, which is obtained by combining the encoding result from the processing unit 135 in the layer one layer below the layer of the processing unit 136 and the decoding result from the processing unit 132, and supplies the decoded result to the combining unit 140. The processing unit 137 decodes the highest layer of the information supplied from the combining unit 140, which is obtained by combining the encoding result from the processing unit 136 in the layer one layer below the layer of the processing unit 137 and the decoding result from the processing unit 131. The processing unit 137 then supplies the decoded result to the VQ-VAE decoder 118 as a cropless image in latent space.
[0130] The combining unit 138 combines the encoding result from the processing unit 133 with the decoding result from the processing unit 134, and supplies the resulting information to the processing unit 135. The combining unit 139 combines the encoding result from the processing unit 132 with the decoding result from the processing unit 135, and supplies the resulting information to the processing unit 136. The combining unit 140 combines the encoding result from the processing unit 131 with the decoding result from the processing unit 136, and supplies the resulting information to the processing unit 137.
[0131] As described above, the image generation unit 117 cross-attentionally connects the latent space input image, the restored image, and the analysis meta-information generated by the latent space image analysis process as prompts. Therefore, compared to when only text data is input as a prompt, the image generation unit 117 can generate images more suitable for computer vision and generate high-quality cropless images of the latent space.
[0132] <Explanation of Crop Interpolation Processing> Fig. 15 is a flowchart illustrating the crop interpolation processing by the image generation unit 117 in Fig. 14. This crop interpolation processing is started, for example, when the input image and restored image of the latent space and analysis meta information are input from the image processing unit 115.
[0133] In step S111 of FIG. 15, the cross-attention I / F 130 of the image generation unit 117 supplies the input image and restored image in the latent space and the analysis meta information supplied from the image processing unit 115 to the processing units 131 to 137, and performs cross-attention connection.
[0134] In step S112, the processing units 131 to 133 and a part of the processing unit 134 perform hierarchical encoding on the latent space image supplied from the VQ-VAE encoder 114. In step S113, a part of the processing unit 134 and the processing units 135 to 137 perform hierarchical decoding on the encoding result of the lowest layer. In step S114, the processing unit 137 outputs the decoding result obtained by the processing of step S113 to the VQ-VAE decoder 118 as a cropless image of the latent space. Then, the crop interpolation process ends.
[0135] As described above, the image generation unit 117 performs cross-attention connection between the input image and the restored image in the latent space obtained as a result of the latent space image analysis process, as well as the analysis meta information. Therefore, the image generation unit 117 can understand the composition of the input image and generate a cropless image in the latent space. As a result, a high-quality cropless image can be generated.
[0136] <Configuration Example of Learning Apparatus> FIG. 16 is a block diagram showing a configuration example of a learning apparatus that performs machine learning of a crop interpolation model.
[0137] 16 includes an acquisition unit 160, a reduction unit 161, an optical noise addition unit 162, a VAE encoder 163, a text encoder 164, an image analysis unit 165, and a cross-attention I / F 166. The learning device 150 also includes an optical noise addition unit 167, a Gaussian noise addition unit 168, a learning unit 169, a denoising loss calculation unit 170, a noise removal unit 171, and a reconstruction loss calculation unit 172.
[0138] The acquisition unit 160 acquires and stores a training set from a server (not shown) or the like. The training data included in this training set is an ideal, high-quality image in pixel space. This training data includes text data in character space corresponding to the image as needed. The acquisition unit 160 sequentially reads out the training data for multiple frames included in the stored training set, one frame at a time. The acquisition unit 160 supplies the pixel space images from the read training data to the reduction unit 161, the optical noise addition unit 167, and the reconstruction loss calculation unit 172. The acquisition unit 160 supplies the text data in character space from the read training data to the text encoder 164 and the training unit 169.
[0139] The reduction unit 161 reduces the resolution of the pixel space image supplied from the acquisition unit 160, similar to the reduction unit 21 in FIG. 2 , and supplies the reduced image to the optical noise addition unit 162. The optical noise addition unit 162 adds a crop region and flicker as optical noise to the image supplied from the reduction unit 161, and supplies the resulting image to the VAE encoder 163. Similar to the VAE encoder 22, the VAE encoder 163 converts the pixel space image supplied from the optical noise addition unit 162 into a latent space image and supplies the resulting image to the image analysis unit 165. Similar to the text encoder 23, the text encoder 164 converts the character space text data supplied from the acquisition unit 160 into latent space text data and supplies the resulting image to the image analysis unit 165.
[0140] The image analysis unit 165 is configured similarly to the image analysis unit 24 in Fig. 13 and performs latent space image analysis processing. The image analysis unit 165 supplies the input image and restored image of the latent space as intermediate or final results of the latent space image analysis processing, as well as analysis meta information, to the cross-attention I / F 166. Note that the analysis meta information of the latent space among the final results of the latent space image analysis processing may be decoded into analysis meta information of the pixel space and output.
[0141] The cross-attention I / F 166 supplies the input image and restored image in the latent space and the analysis meta information supplied from the image analysis unit 165 to the learning unit 169 .
[0142] Similar to the optical noise addition unit 162, the optical noise addition unit 167 adds a crop region and flicker as optical noise to the image supplied from the acquisition unit 160, and supplies the result to the Gaussian noise addition unit 168. The optical noise addition unit 167 also supplies information representing the optical noise to the Gaussian noise addition unit 168.
[0143] The Gaussian noise adding unit 168 adds Gaussian noise to the image supplied from the optical noise adding unit 167, and supplies the result to the learning unit 169 and the noise removing unit 171. The Gaussian noise adding unit 168 also combines the Gaussian noise with the optical noise represented by the information supplied from the optical noise adding unit 167, and supplies the combined result to the denoising loss calculation unit 170 as an expected value of the overall noise.
[0144] The learning unit 169 cross-attentionally connects the input image, restored image, and analysis meta information in the latent space supplied from the cross-attention I / F 166 to a learning model corresponding to the crop interpolation model. Using this learning model, the learning unit 169 generates an inferred value of the overall noise added to the image supplied from the Gaussian noise adding unit 168. Note that text data in the latent space is cross-attentionally connected to this learning model as needed. This text data in the latent space is obtained by encoding the text data in the character space supplied from the acquisition unit 160. The learning unit 169 supplies the generated inferred value of the overall noise to the denoising loss calculation unit 170 and the noise removal unit 171.
[0145] The denoise loss calculation unit 170 calculates a denoise loss value using the expected value of the overall noise supplied from the Gaussian noise addition unit 168 and the inferred value of the overall noise supplied from the learning unit 169. Note that the expected value of the overall noise may be calculated by the denoise loss calculation unit 170 from the difference between the image output from the acquisition unit 160 and the image output from the Gaussian noise addition unit 168. The denoise loss calculation unit 170 uses the calculated denoise loss value to propagate an error to the learning unit 169, thereby performing machine learning of a crop interpolation model.
[0146] The noise removal unit 171 removes the estimated value of overall noise supplied from the learning unit 169 from the image in pixel space to which optical noise and Gaussian noise have been added, supplied from the Gaussian noise addition unit 168, and supplies the result to the reconstruction loss calculation unit 172.
[0147] The reconstruction loss calculation unit 172 calculates a reconstruction loss value using the image in pixel space supplied from the acquisition unit 160 as an expected value and the image in pixel space supplied from the noise removal unit 171 as an inferred value. The reconstruction loss calculation unit 172 uses the reconstruction loss value to propagate an error to the learning unit 169, thereby performing machine learning of a crop interpolation model.
[0148] During error propagation, the denoising loss value and the reconstruction loss value are weighted as necessary.
[0149] As described above, the learning device 150 can learn a crop interpolation model that also removes flicker by adding flicker as optical noise in the optical noise addition unit 167. The learning device 150 can machine-learn a crop interpolation model that also performs electronic motion deblur correction by adding Gaussian noise in the Gaussian noise addition unit 168. Note that the noise added to the image in pixel space in the upstream stage of the learning unit 169 is not limited to flicker or Gaussian noise. The learning device 150 can machine-learn a crop interpolation model that also removes noise by adding noise in the upstream stage of the learning unit 169 using a degradation function of noise that is detrimental to camera signal processing.
[0150] The text data of the character space output from the acquisition unit 160 may also be supplied to the optical noise addition unit 167 and the Gaussian noise addition unit 168. In this case, the optical noise addition unit 167 and the Gaussian noise addition unit 168 add optical noise and Gaussian noise using a degradation function according to the text data, and the learning unit 169 performs machine learning on a crop interpolation model that also removes noise according to the text data.
[0151] As a result, for example, the learning unit 169 can machine-learn a crop interpolation model that also performs electronic motion deblur correction in accordance with text data that instructs the degree of motion blur. The learning unit 169 can machine-learn a crop interpolation model that also performs processing to change the sense of depth by changing lens characteristics such as focal length in accordance with text data that instructs optical sense of depth. The learning unit 169 can machine-learn a crop interpolation model that also performs processing to change the texture of a light source, such as a daytime light source or an evening light source, in accordance with text data that instructs white balance adjustment. As a result, the image generation unit 117 in FIG. 12 can use the thus-machine-learned crop interpolation model to perform crop interpolation processing that also performs processing in accordance with the input text data in the latent space.
[0152] <Explanation of Learning Process> FIG. 17 is a flowchart illustrating the learning process performed by the learning device 150 of FIG.
[0153] 17, the learning device 150 initializes the hyperparameters of the crop interpolation model. In step S132, the learning device 150 initializes the network definition of the crop interpolation model. In step S133, the acquisition unit 160 initializes the currently held training set.
[0154] In step S134, the acquisition unit 160 reads out the training data for frames that have not yet been read out from the training data for each of the multiple frames included in the training set initialized by the processing in step S133. The acquisition unit 160 supplies the pixel space images from the training data to the reduction unit 161, the optical noise addition unit 167, and the reconstruction loss calculation unit 172. The acquisition unit 160 supplies the text data from the read training data to the text encoder 164 or the training unit 169.
[0155] In step S135, the optical noise addition unit 167 adds the crop region and flicker as optical noise to the image in pixel space supplied from the acquisition unit 160, and supplies the result to the Gaussian noise addition unit 168. The optical noise addition unit 167 also supplies information representing the optical noise to the Gaussian noise addition unit 168.
[0156] In step S136, the Gaussian noise adding unit 168 adds Gaussian noise to the image to which optical noise has been added by the processing in step S135, and supplies the result to the learning unit 169 and the noise removing unit 171. The Gaussian noise adding unit 168 also combines the Gaussian noise with the optical noise represented by the information supplied from the optical noise adding unit 167, and supplies the result to the denoising loss calculation unit 170 as an expected value of the overall noise.
[0157] In step S137, the cross-attention I / F 166 supplies the input image and restored image in the latent space, as well as the analysis meta information, to the learning unit 169, and cross-attentionally connects them to the learning model corresponding to the crop interpolation model. The input image and restored image in the latent space, as well as the analysis meta information, are supplied from the image analysis unit 165 to the cross-attention I / F 166 as a result of the image in pixel space being supplied from the acquisition unit 160 to the reduction unit 161.
[0158] In step S138, the learning unit 169 performs forward processing to generate an estimated value of the overall noise using a learning model corresponding to the crop interpolation model cross-attentionally connected in step S137. The learning unit 63 supplies the generated estimated value of the overall noise to the denoising loss calculation unit 170 and the noise removal unit 171.
[0159] In step S139, the noise removal unit 171 removes the estimated value of the overall noise generated in the process of step S138 from the image to which Gaussian noise has been added in the process of step S136, and supplies the result to the reconstruction loss calculation unit 172.
[0160] In step S140, the denoising loss calculation unit 170 calculates a denoising loss value using the expected value of the overall noise supplied from the Gaussian noise addition unit 168 and the inferred value of the overall noise generated by the processing in step S138.
[0161] In step S141, the reconstruction loss calculation unit 172 calculates a reconstruction loss value using the image supplied from the acquisition unit 160 as an expected value and the image from which the estimated value of the overall noise has been removed by the processing of step S139 as an estimated value.
[0162] In step S142, the denoising loss calculation unit 170 uses the denoising loss value calculated in the processing of step S140 or the reconstruction loss value calculated in the processing of step S141 to propagate an error to the learning unit 169. As a result, the crop interpolation model is machine-learned, and the parameters of the crop interpolation model are updated.
[0163] In step S143, the learning device 150 determines whether or not to end the learning process, similar to the process in step S41 of Fig. 10. If it is determined in step S142 that the learning process should not be ended, the process returns to the process in step S134, and the subsequent processes are repeated.
[0164] On the other hand, if it is determined in step S143 that the learning process is to be ended, the process proceeds to step S144. In step S144, the learning unit 63 outputs the parameters of the crop interpolation model updated in the process of the final step S142 as dictionary data. The crop interpolation model to which this dictionary data is set is used in the image generation unit 117 in FIG. 12. After the process of step S144, the learning process ends.
[0165] <Outline of Example of Crop Interpolation Model Learning Method> FIG. 18 is a diagram illustrating an outline of an example of a crop interpolation model learning method.
[0166] As shown in A of Figure 18, the first learning method is a method of machine learning in which a foundation model of a cropless image generation AI is used as a crop interpolation model by performing a learning process using a huge learning set 181. In the first learning method, the learning process is performed using a huge learning set 181, so average machine learning is performed for all scenes. The image generation unit 117 uses the foundation model of the cropless image generation AI generated by the first learning method as a crop interpolation model to generate a cropless image video 182 from a video that is to be subjected to crop interpolation processing.
[0167] As shown in B of Fig. 18 , the second training method is a method of further fine-tuning the foundation model pre-trained by the first training method using a training set 191 of moving images to be subjected to crop interpolation processing. The second training method enables machine learning specialized for scenes of moving images to be subjected to crop interpolation processing. The image generation unit 117 generates a cropless moving image 192 from the moving images to be subjected to crop interpolation processing using the fine-tuned crop interpolation model generated by the second training method.
[0168] As described above, the second learning method performs machine learning specialized for video scenes that are the target of crop interpolation processing. Therefore, the fine-tuned crop interpolation model is more suitable for video scenes that are the target of crop interpolation processing than the foundation model of the AI for generating cropless images. Therefore, the accuracy of the cropless video 192 is higher than the accuracy of the cropless video 182.
[0169] As a fine tuning method in the second learning method, the same method as in the first embodiment can be used.
[0170] <Examples of Image After Electronic Stabilization Processing and Cropless Image> Fig. 19 is a diagram showing an example of an image after electronic stabilization processing by the electronic stabilization processing unit 113 in Fig. 12 and a cropless image generated by the image generation unit 117 using a crop interpolation model learned by machine learning using the first learning method. Fig. 20 is a diagram showing an example of an image after electronic stabilization processing and a cropless image generated by the image generation unit 117 using a crop interpolation model learned by machine learning using the second learning method. Note that in the example of Fig. 20, fine tuning was performed using LoRA.
[0171] 19A and 19B show examples of six images after electronic stabilization processing on the left side, respectively. 19A and 19B show examples of six cropless images that are generated when the six images after electronic stabilization processing are input to the image generation unit 117 on the right side. The same applies to FIG. 20.
[0172] The interpolated cropped areas in the cropless images shown on the right side of Figures 19A and 19B do not contain any landscape information and are filled with plausible low-frequency components. Therefore, the cropped areas appear blurred. Note that the cropped areas may also be filled with significant artifacts.
[0173] In contrast, detailed patterns are restored in the interpolated cropped areas of the cropless images shown on the right side of A and B in Figure 20. This is because fine tuning is performed using a training set of video images that are the subject of crop interpolation processing, and the scenery of the video images is machine-learned as a whole in the crop interpolation model.
[0174] The crop interpolation model training method is selected, for example, depending on the purpose for which the crop interpolation model is used. When the second training method is used to train the crop interpolation model in video editing (post-production), the editing process takes an additional time, equivalent to the time required for fine-tuning, compared to when the first training method is used. Therefore, for example, when a user develops a 60-second video using a crop interpolation model, it is desirable that the fine-tuning time be approximately 30 seconds, which is half the duration of the video. This allows the editing process time when the second training method is used to train the crop interpolation model to be reduced to approximately 1.5 to 2.0 times the editing process time when the first training method is used.
[0175] The image generation unit 117 may perform processing to improve the quality of the latent space image and generate a high-quality latent space image, other than crop interpolation processing that interpolates cropped areas generated by electronic stabilization processing in the rotational direction. For example, the image generation unit 117 may perform occlusion interpolation processing that improves the quality of the latent space image by interpolating occlusion areas generated by electronic stabilization processing in the translational direction, and generates a high-quality latent space image. In this case, the electronic stabilization processing unit 113 performs electronic stabilization processing in the translational direction based on IMU information, and the learning device 150 adds the occlusion areas as optical noise instead of cropped areas.
[0176] The image generator 117 may also perform noise reduction processing to improve the quality of the latent space image by removing noise such as flicker, thereby generating a high-quality image of the latent space. In this case, the electronic stabilization processor 113 is not provided, and the learning device 150 does not add, for example, a crop region as optical noise.
[0177] The image generation unit 117 may further perform electronic motion deblur correction to improve the quality of the latent space image and generate a deblurred image as a high-quality image of the latent space. In this case, the electronic stabilization processing unit 113 is not provided, and the learning device 150 does not add optical noise, for example.
[0178] The image generating unit 117 may perform a combination of two or more of the crop interpolation process, the occlusion interpolation process, the noise reduction process, and the electronic motion deblur correction process.
[0179] The IMU sensor 112 may detect only one of the acceleration and angular velocity required for subsequent processing, and output IMU information indicating only the detected one.
[0180] <Examples of image in pixel space and deblur image> Figure 21 is a diagram showing an example of an image output from the imaging unit 111 when the image generation unit 117 performs electronic motion deblur correction processing, and an example of a deblur image in pixel space obtained as a result of performing electronic motion deblur correction processing on that image.
[0181] 21A shows an example of six images output from the imaging unit 111, and FIG. 21B shows an example of six deblurred images obtained as a result of performing electronic motion deblur correction processing on each of the six pixel space images. As shown in FIG. 21B, the image generating unit 117 can generate accurate deblurred images by the electronic motion deblur correction processing.
[0182] 3. Third Embodiment Configuration Example of Imaging System FIG. 22 is a block diagram showing a configuration example of an imaging system according to a third embodiment to which the present technology is applied.
[0183] In the imaging system 210 of Fig. 22, the parts corresponding to those in the imaging system 110 of Fig. 12 are denoted by the same reference numerals. Therefore, the description of those parts will be omitted as appropriate, and the following description will focus on the parts that are different from the imaging system 110.
[0184] 22 includes an input unit 12, an imaging unit 111, an IMU sensor 112, an image processing unit 115, a text encoder 116, an IMU analysis unit 212, and decoders 213 to 217. The imaging system 210 analyzes IMU information and performs IMU analysis processing to generate IMU analysis information (detection information analysis information) that represents the results.
[0185] Specifically, the IMU analysis unit 212 (detection information analysis unit) cross-attentionally connects the input image and restored image in the latent space, as well as the analysis meta information, from the image processing unit 115 to an IMU analysis model (detection information analysis model), which is a learning model for the AI used to analyze the IMU information. The IMU analysis unit 212 also cross-attentionally connects text data in the latent space supplied from the text encoder 116 to the IMU analysis model as needed. This text data includes text data expressing instructions related to electronic stabilization processing or optical stabilization processing based on the content of the image, such as not to eliminate too much realism. Using this IMU analysis model, the IMU analysis unit 212 analyzes the IMU information in the sensor information space output from the IMU sensor 112 and generates IMU analysis information in the latent space.
[0186] The IMU analysis information is composed of local motion (LMV) information, rotation / translation information, calibrated IMU information (Gyro), and STB control information (Crtl). The local motion information is information that represents the movement of a subject within an image in pixel space. The rotation / translation information is composed of rotation information that represents movement in the rotational direction and translation information that represents movement in the translational direction, both of which are part of global motion (GMV) information that represents the movement of the imaging unit 111. The rotation information is used for electronic stabilization processing in the rotational direction, and the translation information is used for electronic stabilization processing in the translational direction.
[0187] The calibrated IMU information is IMU information from which a bias specific to the IMU sensor 112 has been removed, i.e., is obtained by comparing the angular velocity represented by the IMU information with the angular velocity resulting from analysis of the image output from the imaging unit 111. The calibrated IMU information is used for optical stabilization processing and the like in driving the lens and image sensor in the imaging unit 111. The STB control information is attitude control information that controls the attitude of the imaging system 210 taking into account image edge contact during electronic stabilization processing, and is used for electronic stabilization processing and optical stabilization processing for panning functions and the like.
[0188] The IMU analysis unit 212 supplies the generated local motion information, rotation information, translation information, calibrated IMU information, and STB control information of the latent space to the decoders 213 to 217, respectively.
[0189] The decoder 213 (detection information analysis information decoding unit) decodes the STB control information in the latent space supplied from the IMU analysis unit 212, converts it into STB control information in the sensor information space (detection information space), and outputs it. The decoder 214 (detection information analysis information decoding unit) decodes the calibrated IMU information in the latent space supplied from the IMU analysis unit 212, converts it into calibrated IMU information in the sensor information space, and outputs it. The decoder 215 (detection information analysis information decoding unit) decodes the rotation information in the latent space supplied from the IMU analysis unit 212, converts it into rotation information in the sensor information space, and outputs it. The decoder 216 (detection information analysis information decoding unit) decodes the translation information in the latent space supplied from the IMU analysis unit 212, converts it into translation information in the sensor information space, and outputs it. The decoder 217 (detection information analysis information decoding unit) decodes the local motion information in the latent space supplied from the IMU analysis unit 212, converts it into local motion information in the sensor information space, and outputs it.
[0190] In this way, the imaging system 210 generates IMU analysis information in the sensor information space. Therefore, a downstream processing unit (not shown) uses this IMU analysis information to perform electronic stabilization processing or optical stabilization processing on the moving images captured by the imaging unit 111, thereby achieving high-quality electronic image stabilization or optical image stabilization.
[0191] <Configuration Example of IMU Analysis Unit> FIG. 23 is a block diagram showing a configuration example of the IMU analysis unit 212 in FIG. 22 .
[0192] The IMU analysis unit 212 in Fig. 23 has a U-Net structure. Specifically, the IMU analysis unit 212 includes a cross-attention I / F 230, an IMU encoder 231, encoders 232 to 234, a GMT / LMT separation decoder 235, and a combination unit 236. The IMU analysis unit 212 also includes a rotation / translation separation decoder 237, a combination unit 238, a bias separation decoder 239, a combination unit 240, and a camerawork decoder 241.
[0193] The cross-attention I / F 230 supplies the input image and restored image in the latent space and the analysis meta information supplied from the image processing unit 115 in Fig. 22 to the IMU encoder 231 and the encoders 231 to 234. The cross-attention I / F 230 also supplies the input image and restored image in the latent space and the analysis meta information to the GMT / LMT separation decoder 235, the rotation / translation separation decoder 237, the bias separation decoder 239, and the camerawork decoder 241.
[0194] The IMU encoder 231 and encoders 232 to 234 are connected hierarchically, and perform four-layer hierarchical encoding on the IMU information supplied from the IMU sensor 112 in Fig. 22. The GMT / LMT separation decoder 235, rotation / translation separation decoder 237, bias separation decoder 239, and camerawork decoder 241 are connected hierarchically, and perform four-layer hierarchical decoding on the encoding result of the lowest layer of the four layers.
[0195] Specifically, the IMU encoder 231 and the encoders 232 to 234 are cross-attentionally connected to the input image and restored image in the latent space, as well as the analysis meta information, supplied from the cross-attention I / F 230. The GMT / LMT separation decoder 235 and the rotation / translation separation decoder 237 are cross-attentionally connected to the input image and restored image in the latent space, as well as the analysis meta information, supplied from the cross-attention I / F 230. The bias separation decoder 239 and the camerawork decoder 241 are cross-attentionally connected to the input image and restored image in the latent space, as well as the analysis meta information, supplied from the cross-attention I / F 230.
[0196] The IMU encoder 231 is an encoder that encodes the top layer of the four layers. The IMU encoder 231 is skip-connected to the camerawork decoder 241 via the combining unit 240. The IMU encoder 231 converts the IMU information in the sensor information space supplied from the IMU sensor 112 into IMU information in the latent space and supplies the converted information to the encoder 232 and the combining unit 240.
[0197] The encoder 232 is an encoder that performs encoding at the second highest layer, and includes a residual combination block and a self-attention mechanism block. The encoder 232 is skip-connected to the bias separation decoder 239 via the combination unit 238. The encoder 232 encodes the encoding result provided by the encoder 231 and provides the encoding result to the encoder 233 and the combination unit 238.
[0198] The encoder 233 is an encoder that performs encoding at the third layer from the top, and includes a residual combination block and a self-attention mechanism block. The encoder 233 is skip-connected to the rotation / translation separation decoder 237 via the combination unit 236. The encoder 233 encodes the encoding result provided by the encoder 232 and provides the encoding result to the encoder 234 and the combination unit 236.
[0199] The encoder 234 is an encoder that performs encoding at the lowest layer and includes a residual combination block and a self-attention mechanism block. The encoder 234 is skip-connected to the GMT / LMT separation decoder 235. The encoder 233 encodes the encoding result provided by the encoder 233 and provides the encoding result to the GMT / LMT separation decoder 235.
[0200] The GMT / LMT separation decoder 235 is a decoder that performs decoding on the lowest layer of the four layers, and includes a residual combination block and a self-attention mechanism block. The GMT / LMT separation decoder 235 separates the motion analysis information in the latent space into global motion analysis information and local motion information in the latent space by decoding the encoding result supplied from the encoder 234. The GMT / LMT separation decoder 235 supplies the local motion information in the latent space to the decoder 217 in FIG. 22 and supplies the decoded result to the combination unit 236.
[0201] The combining unit 236 combines the encoding result supplied from the encoder 233 with the decoding result supplied from the GMT / LMT separation decoder 235 , and supplies the resulting information to a rotation / translation separation decoder 237 .
[0202] The rotation / translation separation decoder 237 is a decoder that performs decoding at the third layer from the top, and includes a residual combination block and a self-attention mechanism block. The rotation / translation separation decoder 237 decodes the information provided by the combination unit 236 to further separate the global motion analysis information in the latent space into translation information and rotation information in the latent space. The rotation / translation separation decoder 237 provides the rotation information in the latent space to the decoder 215 in FIG. 22 and the translation information in the latent space to the decoder 216. The rotation / translation separation decoder 237 also provides the decoded results to the combination unit 238.
[0203] The combiner 238 combines the encoding result supplied from the encoder 231 with the decoding result supplied from the rotation / translation separation decoder 237 , and supplies the resulting information to a bias separation decoder 239 .
[0204] The bias separation decoder 239 is a decoder that performs decoding at the second highest layer and includes a residual combination block and a self-attention mechanism block. The bias separation decoder 239 generates calibrated IMU information in the latent space by decoding the information provided by the combination unit 238. The bias separation decoder 239 provides the calibrated IMU information in the latent space to the decoder 214 in FIG. 22 and provides the decoded result to the combination unit 240.
[0205] The combining unit 240 combines the IMU information in the latent space supplied from the IMU encoder 231 with the decoding result supplied from the bias separation decoder 239, and supplies the resulting information to the camerawork decoder 241.
[0206] The camerawork decoder 241 is a decoder that performs decoding at the top layer and includes a residual combination block and a self-attention mechanism block. The camerawork decoder 241 decodes the information supplied from the combination unit 240 and generates STB control information in the latent space. The camerawork decoder 241 supplies the STB control information in the latent space to the decoder 213 in FIG. 22.
[0207] As described above, in the IMU analysis unit 212, each layer has a function of generating, from top to bottom, STB control information, calibrated IMU information, rotation / translation information, and local motion information in the latent space. In the IMU analysis unit 212, the input image and restored image in the latent space generated by the latent space image analysis process, as well as the analysis meta information, are cross-attention-connected, so the IMU analysis unit 212 can generate IMU analysis information after understanding the structure of the image.
[0208] <Explanation of IMU Analysis Processing> Fig. 24 is a flowchart illustrating the IMU analysis processing by the imaging system 210 of Fig. 22. This IMU analysis processing starts, for example, when an input image and a restored image in the latent space and analysis meta information are input from the image processing unit 115.
[0209] In step S211 of Fig. 24, the cross-attention I / F 230 supplies the input image and restored image in the latent space and the analysis meta information from the image processing unit 115 to the IMU encoder 231 and the encoders 232 to 234, and performs cross-attention connection. The cross-attention I / F 230 also supplies the input image and restored image in the latent space and the analysis meta information to the GMT / LMT separation decoder 235 and the rotation / translation separation decoder 237, and performs cross-attention connection. The cross-attention I / F 230 also supplies the input image and restored image in the latent space and the analysis meta information to the bias separation decoder 239 and the camerawork decoder 241, and performs cross-attention connection.
[0210] In step S212, the IMU analysis unit 212 generates IMU analysis information in the latent space from the IMU information in the sensor information space supplied from the IMU sensor 112. The IMU analysis unit 212 supplies local motion information in the latent space, which is included in the IMU analysis information in the latent space, to the decoder 217, translation information to the decoder 216, and rotation information to the decoder 215. The IMU analysis unit 212 also supplies calibrated IMU information in the latent space, which is included in the IMU analysis information in the latent space, to the decoder 214, and supplies STB control information to the decoder 213.
[0211] In step S213, the decoder 217 converts the local motion information in the latent space generated by the processing of step S212 into local motion information in the sensor information space and outputs it. In step S214, the decoder 216 converts the translation information in the latent space generated by the processing of step S212 into translation information in the sensor information space and outputs it. In step S215, the decoder 215 converts the rotation information in the latent space generated by the processing of step S212 into rotation information in the sensor information space and outputs it.
[0212] In step S216, the decoder 214 converts the calibrated IMU information in the latent space generated by the processing of step S212 into calibrated IMU information in the sensor information space and outputs it. In step S217, the decoder 213 converts the STB control information in the latent space generated by the processing of step S212 into STB control information in the sensor information space and outputs it. Then, the IMU analysis processing ends.
[0213] <Configuration Example of Learning Apparatus> FIG. 25 is a block diagram showing a configuration example of a learning apparatus that performs machine learning to learn an IMU analysis model.
[0214] In the learning device 250 in Fig. 25, parts corresponding to those in the learning device 150 in Fig. 16 are denoted by the same reference numerals. Therefore, the description of those parts will be omitted as appropriate, and the following description will focus on the parts that are different from the learning device 150.
[0215] 25 includes a reduction unit 161, an optical noise addition unit 162, a VAE encoder 163, a text encoder 164, an image analysis unit 165, and a cross-attention I / F 166. The learning device 250 also includes an acquisition unit 260, an offset addition unit 261, a learning unit 262, and loss calculation units 263 to 266.
[0216] The acquisition unit 260 acquires and stores a training set from a server (not shown). The training data included in this training set is composed of ideal, high-quality pixel space images, IMU information in sensor information space corresponding to the images, and expected values of some of the IMU analysis information. The training data also includes text data in character space corresponding to the images, as needed.
[0217] The acquisition unit 260 sequentially reads out, frame by frame, the learning data for multiple frames included in the held learning set. The acquisition unit 260 supplies the pixel space image of the read learning data to the reduction unit 161, and supplies the IMU information to the offset addition unit 261. The acquisition unit 260 supplies the expected value of the local motion information of the read learning data to the loss calculation unit 264, supplies the expected value of the rotation / translation information to the loss calculation unit 265, and supplies the expected value of the STB control information to the loss calculation unit 266. The acquisition unit 260 supplies the text data of the character space of the read learning data to the text encoder 164 and the learning unit 262.
[0218] The offset adding unit 261 adds noise such as an offset to the IMU information supplied from the acquisition unit 260 and supplies the information to the learning unit 262 and loss calculation unit 263.
[0219] The learning unit 262 cross-attentionally connects the input image, restored image, and analysis meta information in the latent space supplied from the cross-attention I / F 166 to the IMU analysis model or a learning model corresponding to the IMU analysis model. The learning unit 262 uses the IMU analysis model to generate IMU analysis information in the sensor information space from the IMU information supplied from the offset adding unit 261. Alternatively, the learning unit 262 uses the learning model corresponding to the IMU analysis model to generate, from the IMU information, an inferred value of the offset added to the IMU information and the IMU analysis information in the sensor information space other than the calibrated IMU information.
[0220] In the following, the process in which the learning unit 262 generates IMU analysis information in the sensor information space from IMU information, or an inferred value of the offset and the IMU analysis information in the sensor information space other than the calibrated IMU information, is referred to as IMU analysis forward processing.
[0221] The learning unit 262 supplies the calibrated IMU information or offset inference value of the generated sensor information space to the loss calculation unit 263, and supplies the local motion information to the loss calculation unit 264. The learning unit 262 supplies the rotational / translation information of the generated sensor information space to the loss calculation unit 265, and supplies the STB control information to the loss calculation unit 266. Note that, like the learning unit 263, time embedding information is input to the learning unit 262.
[0222] The loss calculation unit 263 calculates a reconstruction loss value using the IMU information without offset supplied from the acquisition unit 260 as an expected value (teacher data) and the calibrated IMU information supplied from the learning unit 262 as an inferred value. Alternatively, the loss calculation unit 263 generates an expected offset value, which is the difference between the IMU information without offset and the IMU information with offset supplied from the offset addition unit 261. The loss calculation unit 263 then calculates a denoising loss value using the expected offset value and the inferred offset value supplied from the learning unit 262. The loss calculation unit 263 performs machine learning of the IMU analysis model by propagating an error to the learning unit 262 using the calculated reconstruction loss value or denoising loss value. This updates the parameters of the IMU analysis model.
[0223] When the reconstruction loss value is error-propagated, the learning unit 262 can machine-learn an IMU analysis model that generates calibrated IMU information that is faithful to the input IMU information. Therefore, the IMU analysis process using the IMU analysis model learned in this way is primarily intended to correct the input IMU information.
[0224] On the other hand, when the denoising loss value is error propagated, it is possible to machine-learn an IMU analysis model that can generate calibrated IMU information including information not present in the IMU information input as a prompt, from the IMU information, etc. Therefore, by performing IMU analysis processing using the IMU analysis model machine-learned in this way, it is possible to perform image stabilization effect processing, and to generate STB control information for performing image stabilization without losing the sense of realism, taking into account the image composition and vibration level.
[0225] The loss calculation unit 264 calculates a reconstruction loss value using the expected value of the local motion information supplied from the acquisition unit 260 and (the inferred value of) the local motion information supplied from the learning unit 262. The loss calculation unit 264 performs machine learning of the IMU analytical model by using the reconstruction loss value to propagate an error to the learning unit 262. As a result, the parameters of the IMU analytical model are updated.
[0226] The loss calculation unit 265 calculates a reconstruction loss value using the expected value of the rotation / translation information supplied from the acquisition unit 260 and (the inferred value of) the rotation / translation information supplied from the learning unit 262. The loss calculation unit 265 performs machine learning of the IMU analytical model by using the reconstruction loss value to propagate an error to the learning unit 262. As a result, the parameters of the IMU analytical model are updated.
[0227] The loss calculation unit 266 calculates a reconstruction loss value using the expected value of the STB control information supplied from the acquisition unit 260 and (the inferred value of) the STB control information supplied from the learning unit 262. The loss calculation unit 266 performs machine learning of the IMU analytical model by error-propagating the reconstruction loss value to the learning unit 262. As a result, the parameters of the IMU analytical model are updated.
[0228] The loss calculation units 263 to 266 weight the calculation of the reconstruction loss value or the denoising loss value as necessary.
[0229] <Explanation of Learning Process> FIG. 26 is a flowchart illustrating the learning process performed by the learning device 250 of FIG.
[0230] 26 , in step S231, the learning device 250 initializes the hyperparameters of the IMU analytical model. In step S232, the learning device 250 initializes the network definition of the IMU analytical model. In step S233, the acquisition unit 260 initializes the currently held learning set.
[0231] In step S234, the acquisition unit 260 reads out the learning data of frames that have not yet been read out from the learning data of each of the multiple frames included in the learning set initialized by the processing of step S233. The acquisition unit 260 supplies the pixel space image of the learning data to the reduction unit 161, and supplies the IMU information to the offset addition unit 261 and the loss calculation unit 263. The acquisition unit 260 supplies the expected value of the local motion information in the sensor information space of the read learning data to the loss calculation unit 264, supplies the expected value of the rotational / translation information to the loss calculation unit 265, and supplies the expected value of the STB control information to the loss calculation unit 266. The acquisition unit 260 supplies the text data of the read learning data to the image analysis unit 165 and the learning unit 262.
[0232] In step S235, the offset adding unit 261 adds noise such as an offset to the IMU information supplied from the acquisition unit 260 and supplies the IMU information to the learning unit 262 and the loss calculation unit 263.
[0233] In step S236, the cross-attention I / F 166 supplies the input image and restored image in the latent space, as well as the analysis meta-information, to the learning unit 262, and cross-attentionally connects them to the IMU analysis model or a learning model corresponding to the IMU analysis model. The input image and restored image in the latent space, as well as the analysis meta-information, are supplied from the image analysis unit 165 to the cross-attention I / F 166 as a result of the image in pixel space being supplied from the acquisition unit 160 to the reduction unit 161.
[0234] In step S237, the learning unit 262 performs IMU analysis forward processing. The learning unit 262 supplies the resulting calibrated IMU information or offset inference value, local motion information, rotation / translation information, and STB control information to the loss calculation units 263 to 266, respectively.
[0235] In step S238, the loss calculation unit 263 calculates a reconstruction loss value using the IMU information supplied from the acquisition unit 260 and the calibrated IMU information obtained as a result of the processing in step S237. Alternatively, the loss calculation unit 263 calculates a denoising loss value using an expected value of the offset and an inferred value of the offset obtained as a result of the processing in step S237. Note that the expected value of the offset is the difference between the IMU information supplied from the acquisition unit 260 and the IMU information to which the offset has been added in the processing in step S235. The loss calculation units 264 to 266 each calculate a reconstruction loss value using the expected value supplied from the acquisition unit 260 and the inferred value obtained as a result of the processing in step S237.
[0236] In step S239, the loss calculation units 263 to 266 use the reconstruction loss value or the denoising loss value calculated in the process of step S238 to propagate errors to the learning unit 262, thereby performing machine learning on the IMU analytical model. As a result, the parameters of the IMU analytical model are updated.
[0237] In step S240, the learning device 250 determines whether or not to end the learning process, similar to the process in step S41 of Fig. 10. If it is determined in step S240 that the learning process should not be ended, the process returns to the process in step S234, and the subsequent processes are repeated.
[0238] On the other hand, if it is determined in step S240 that the learning process is to end, the process proceeds to step S241. In step S241, the learning unit 63 outputs the parameters of the IMU analysis model updated in the final process of step S239 as dictionary data. The IMU analysis model in which this dictionary data is set is used by the IMU analysis unit 212 in FIG. 22. After the process of step S241, the learning process ends.
[0239] <Outline of Example of Learning Method of IMU Analysis Model> FIG. 27 is a diagram illustrating an outline of an example of a learning method of an IMU analysis model.
[0240] 27A, the first learning method is a method of performing machine learning on a foundation model of an IMU information analysis AI as an IMU analysis model by performing a learning process using a huge learning set 281. In the first learning method, the learning process is performed using a huge learning set 281, so average machine learning is performed for all scenes. The IMU analysis unit 212 uses the foundation model of the IMU information analysis AI generated by the first learning method as the IMU analysis model to generate IMU analysis information 282 from the IMU information that is the target of the IMU analysis process.
[0241] As shown in B of FIG. 27 , the second learning method is a method of further fine-tuning the foundation model pre-trained by the first learning method using a learning set 291 of IMU information that is the target of IMU analysis processing. The second learning method enables machine learning specialized for the IMU information that is the target of IMU analysis processing. The IMU analysis unit 212 generates IMU analysis information 292 from the IMU information that is the target of IMU analysis processing using the fine-tuned IMU analysis model generated by the second learning method.
[0242] As described above, the second learning method performs machine learning specialized for the IMU information that is the target of the IMU analysis process. Therefore, the fine-tuned IMU analysis model is more suitable for the IMU information that is the target of the IMU analysis process than the foundation model of the IMU information analysis AI. Therefore, the accuracy of the IMU analysis information 292 is higher than the accuracy of the IMU analysis information 282.
[0243] As a fine tuning method in the second learning method, the same method as in the first embodiment can be used.
[0244] 4. Fourth Embodiment Configuration Example of Imaging System FIG. 28 is a block diagram showing a configuration example of an imaging system according to a fourth embodiment to which the present technology is applied.
[0245] In the imaging system 310 of Fig. 28, parts corresponding to those in the imaging system 110 of Fig. 12 and the imaging system 210 of Fig. 22 are denoted by the same reference numerals. Therefore, the description of those parts will be omitted as appropriate, and the description will focus on parts that are different from the imaging systems 110 and 210.
[0246] 28 includes an input unit 12, an imaging unit 111, an IMU sensor 112, a VQ-VAE encoder 114, a text encoder 116, a VQ-VAE decoder 118, and decoders 215 and 216. The imaging system 310 also includes a rotational stabilization processing unit 311, an IMU analysis unit 312, an image processing unit 313, a depth encoder 314, a translational stabilization processing unit 315, a text encoder 316, and an image generation unit 317. The imaging system 310 performs electronic stabilization processing in the rotational and translational directions on moving images captured by the imaging unit 111, and generates an occlusion-less image in which an occlusion region generated by the electronic stabilization processing in the translational direction is interpolated.
[0247] Specifically, the rotational stabilization processing unit 311 performs electronic stabilization processing in the rotational direction on the image in pixel space supplied from the imaging unit 111, based on the rotation information in sensor information space supplied from the decoder 215. The rotational stabilization processing unit 311 supplies the resulting image in pixel space to the translational stabilization processing unit 315, and also supplies it to the image processing unit 313 as an input image.
[0248] The IMU analysis unit 312 (detection information analysis unit) cross-attentionally connects the motion analysis information in the latent space supplied from the image processing unit 313 to the IMU analysis model. The IMU analysis unit 312 also cross-attentionally connects text data in the latent space input from the text encoder 116 to the IMU analysis model as needed. This text data includes text data that indicates instructions for electronic stabilization processing or optical stabilization processing based on the content of the image, such as creating a sense of floating. The IMU analysis unit 312 uses this IMU analysis model to generate IMU analysis information in the latent space from the IMU information in the sensor information space supplied from the IMU sensor 112.
[0249] For example, when text data instructing a floating sensation is cross-attentionally connected, the IMU analysis unit 312 of the imaging system 310, which is a handheld camera, generates translational information for the production camera in the latent space. The translational information for the production camera is translational information for performing translational stabilization processing to smoothly remove translational vibrations of the production camera mounted on a rail. The translational stabilization processing unit 315 uses this translational information for the production camera to perform translational stabilization processing on the image supplied from the rotational stabilization processing unit 311, thereby generating an image with a floating sensation.
[0250] The IMU analysis unit 312 supplies the rotation information of the generated IMU analysis information of the latent space to the decoder 215, and supplies the translation information to the decoder 216. The rotation information of the latent space supplied to the decoder 215 is converted by the decoder 215 into rotation information of the sensor information space and supplied to the rotation stabilization processing unit 311. The translation information of the latent space supplied to the decoder 216 is converted by the decoder 216 into translation information of the sensor information space and supplied to the translation stabilization processing unit 315.
[0251] The image processing unit 313 receives text data from the input unit 12. The image processing unit 13 performs latent space image analysis processing on the input image supplied from the rotation stabilization processing unit 311. The image processing unit 313 supplies latent space depth analysis information obtained as a result of the latent space image analysis processing to the depth encoder 314 and the image generation unit 317, and supplies motion analysis information to the IMU analysis unit 312.
[0252] The depth encoder 314 converts the depth analysis information in the latent space supplied from the image processing unit 313 into depth analysis information in the pixel space, and supplies it to the translational stabilization processing unit 315 .
[0253] The translational stabilization processor 315 performs electronic stabilization processing in the translational direction on the image supplied from the rotational stabilization processor 311, based on the depth analysis information supplied from the depth encoder 314 and the translation information supplied from the decoder 216. The electronic stabilization processing in the translational direction is processing that warps each pixel of the image so as to remove vibrations in the translational direction of each pixel. In this processing, for example, a smaller warp amount is set for pixels with a large depth, i.e., pixels where the subject is farther away, compared to pixels with a small depth.
[0254] In the image after translational electronic stabilization, camera shake in the translational direction is corrected, but occlusion regions occur. An occlusion region is a black region that does not contain image information, or a region to which a flag indicating an occlusion region is assigned in the alpha channel. The translational stabilization processing unit 315 supplies the pixel space image after translational electronic stabilization to the VQ-VAE encoder 114. As a result, the VQ-VAE encoder 114 converts the pixel space image into a latent space image and supplies it to the image generation unit 317.
[0255] The text encoder 316 converts the text data in the character space supplied from the input unit 12 into text data in the latent space and supplies it to the image generation unit 317 .
[0256] The image generation unit 317 cross-attentionally connects the depth analysis information of the latent space supplied from the image processing unit 313 to the occlusion interpolation model. The occlusion interpolation model is a learning model of an AI for generating occlusion-less images, having a U-Net structure. The image generation unit 317 also cross-attentionally connects the text data of the latent space supplied from the text encoder 316 to the occlusion interpolation model as needed.
[0257] The image generation unit 317 uses this occlusion interpolation model to perform occlusion interpolation processing on the image in latent space supplied from the VQ-VAE encoder 114. The occlusion interpolation processing is processing for improving the quality of the image in latent space by interpolating the occlusion region, and generating an occlusion-less image as a high-quality image in the latent space.
[0258] The image generation unit 317 supplies the generated occlusion-free image in the latent space to the VQ-VAE decoder 118. As a result, the VQ-VAE decoder 118 converts the occlusion-free image in the latent space into an occlusion-free image in the pixel space and outputs it.
[0259] The imaging system 310 may be provided with only one of the rotational stabilization processing unit 311 and the translational stabilization processing unit 315 .
[0260] The imaging system 110 (310) may include a depth sensor, an infrared sensor, or the like to realize an autofocus function of the imaging unit 111. In this case, for example, the image generation unit 117 cross-attentionally connects the detection information of the depth sensor or the infrared sensor to the crop interpolation model (occlusion interpolation model).
[0261] As described above, the imaging system 310 uses the moving images captured by the imaging unit 111 and the IMU information generated by the IMU sensor 112 to perform processing to generate an occlusion-less image as multimodal AI processing.
[0262] 5. Fifth Embodiment Configuration Example of Image Processing System FIG. 29 is a block diagram showing a configuration example of an image processing system including an imaging device that is a fifth embodiment of an imaging system to which the present technology is applied.
[0263] In the image processing system 410 of Fig. 29, the same reference numerals are used to designate parts corresponding to those in the imaging system 110 of Fig. 12. Therefore, the description of those parts will be omitted as appropriate, and the following description will focus on parts that are different from the imaging system 110.
[0264] The image processing system 410 in Fig. 29 is configured with an imaging device 411 and an information processing device 412. The imaging device 411 generates analysis meta information, a cropless image, and IMU analysis information for each frame of image from a moving image captured by the imaging unit 111. The information processing device 412 generates a new image using the analysis meta information and the cropless image.
[0265] Specifically, the imaging device 411 is composed of an imaging unit 111 , an IMU sensor 112 , an image processing unit 421 , an image generation unit 422 , and an IMU analysis unit 423 .
[0266] The image processing unit 421 is composed of the electronic stabilization processing unit 113, the image processing unit 115 without the text encoder 23, and decoders 26 to 28. The image processing unit 421 performs electronic stabilization processing in the rotational direction on the pixel space image supplied from the imaging unit 111 using IMU information supplied from the IMU sensor 112, and performs latent space image analysis processing using the resulting image as an input image. The image processing unit 421 supplies the latent space input image and restored image as intermediate or final results of the latent space image analysis processing, as well as analysis meta information, to the image generation unit 422 and the IMU analysis unit 423. The image processing unit 421 converts the analysis meta information from the final result of the latent space image analysis processing into pixel space analysis meta information and supplies it to the output unit 424.
[0267] The image generation unit 422 is composed of the electronic stabilization processing unit 113, the VQ-VAE encoder 114, the image generation unit 117, and the VQ-VAE decoder 118. The image generation unit 422 performs electronic stabilization processing in the rotational direction on the pixel space image supplied from the imaging unit 111 using the IMU information supplied from the IMU sensor 112, and converts the resulting pixel space image into a latent space image. The image generation unit 422 performs crop interpolation processing on the latent space image using a crop interpolation model in which the latent space input image, restored image, and analysis meta information supplied from the image processing unit 421 are cross-attentionally connected. The image generation unit 422 converts the resulting latent space cropless image into a pixel space cropless image and supplies it to the output unit 424.
[0268] The IMU analysis unit 423 is composed of the IMU analysis unit 212 and decoders 213 to 217. The IMU analysis unit 423 generates IMU analysis information in the sensor information space from the IMU information from the IMU sensor 112 using an IMU analysis model in which the input image and restored image in the latent space and the analysis meta information supplied from the image processing unit 421 are cross-attentionally connected. The IMU analysis unit 423 outputs the IMU analysis information in the sensor information space as needed. This IMU analysis information is used, for example, to control the image capture device 411.
[0269] The output unit 424 associates the analysis meta information supplied for each frame from the image processing unit 421 with the cropless image supplied for each frame from the image generating unit 422 in a frame synchronization manner, and outputs the result to the information processing device 412 .
[0270] The information processing device 412 is an editing device, a content production device, or the like, that includes an image generation unit 431. The image generation unit 431 cross-attentionally connects the analysis meta information output from the output unit 424 and the text data input by the user to a learning model of an image generation AI, such as a diffusion model. The image generation unit 431 uses this learning model to generate a new image from the cropless image supplied from the output unit 424.
[0271] As described above, the imaging device 411 supplies not only cropless images but also analytical meta information to the information processing device 412. Therefore, the information processing device 412 can cross-attentionally connect the analytical meta information to the learning model of the image generation AI, thereby improving the performance of the learning model.
[0272] In addition, the imaging device 411 may not be equipped with an image generation unit 422 or an IMU analysis unit 423, and may instead associate the restored image of the pixel space of each frame generated by the image processing unit 421 with the analysis meta information and output it from the output unit 424 to the information processing device 412.
[0273] 6. Sixth Embodiment Configuration Example of Imaging System FIG. 30 is a block diagram showing a configuration example of an imaging system according to a sixth embodiment to which the present technology is applied.
[0274] In the imaging system 510 of Fig. 30, parts corresponding to those in the imaging system 110 of Fig. 12 are assigned the same reference numerals. Therefore, the description of those parts will be omitted as appropriate, and the description will focus on parts that differ from the imaging system 110. The imaging system 510 of Fig. 30 differs from the imaging system 110 in that it does not include the VQ-VAE encoder 114 and the VQ-VAE decoder 118, and that it includes an image generation unit 511 instead of the image generation unit 117. Other than that, it is configured in the same way as the imaging system 110.
[0275] Specifically, the image generation unit 511 attentionally connects the input image and restored image in the latent space and the analysis meta information supplied from the image processing unit 115 to a lightweight crop interpolation model, which is a lightweight crop interpolation model. The image generation unit 511 also attentionally connects text data in the latent space supplied from the text encoder 116 to the lightweight crop interpolation model as needed. This text data includes text data representing instructions for generating a cropless image, such as the optical depth of the cropless image, the degree of motion blur, and white balance adjustment based on a more natural sun light source estimation.
[0276] The image generation unit 511 uses this light crop interpolation model to perform light crop interpolation processing to generate a cropless image in pixel space from the image in pixel space supplied from the electronic stabilization processing unit 113. Note that the resolution of the image is the same before and after the light crop interpolation processing. For example, if the resolution of the image output from the imaging unit 111 is 1280 × 960 pixels, the resolution of the cropless image generated by the light crop interpolation processing is also 1280 × 960 pixels. The image generation unit 511 outputs the cropless image in pixel space generated by the light crop interpolation processing.
[0277] <Configuration Example of Image Generation Unit> FIG. 31 is a block diagram showing a configuration example of the image generation unit 511 in FIG.
[0278] The image generation unit 511 in Figure 31 is composed of an attention I / F 530, a pyramid generation unit 531, AE encoders 532 to 534, a self-attention mechanism 535, a raster conversion unit 536, processing units 537 to 539, combining units 540 to 542, and an AE decoder 543.
[0279] The attention I / F 530 supplies the input image and restored image in the latent space and the analysis meta information supplied from the image processing unit 115 in FIG. 30 to the self-attention mechanism 535 .
[0280] The pyramiding unit 531 performs pyramiding by repeatedly downsampling the image supplied from the electronic stabilization processing unit 113 in Fig. 30 a plurality of times, thereby generating a pyramid image made up of images with three resolutions. The highest resolution of these three resolutions is the same as the resolution of the image before downsampling, i.e., the resolution of the image supplied from the electronic stabilization processing unit 113.
[0281] The pyramiding unit 531 divides each image constituting the pyramid image into processing unit blocks for the light crop interpolation process. The size of the processing unit blocks corresponds to, for example, 3x3 pixels of the image before downsampling. The pyramiding unit 531 outputs the processing unit blocks of each image constituting the pyramid image to each of the AE encoders 532 to 534 as target blocks in raster scan order.
[0282] The AE encoders 532 to 534 (pyramid image encoding units) each convert a target block in pixel space supplied from the pyramiding unit 531 into a target block in latent space. Because these AE encoders 532 to 534 perform conversion on the target block, they can perform conversion with less computational cost than the VQ-VAE encoder 114 in Fig. 12, which performs conversion on the entire image.
[0283] The self-attention mechanism 535 generates attention information of a two-dimensional latent space corresponding to the two-dimensional pixel arrangement from the input image and restored image of the latent space and the analysis meta information supplied from the attention I / F 530, and supplies it to the raster conversion unit 536.
[0284] The raster converter 536 (divider) converts the format of the attention information in the two-dimensional latent space supplied from the self-attention mechanism 535 into attention information in the one-dimensional latent space arranged in raster scan order. The raster converter 536 divides the one-dimensional attention information into processing unit blocks and outputs it to the combiner 540 in raster scan order as attention information for the target block.
[0285] The processing units 537 to 539 are the processing units for the bottom layer, the second-highest layer, and the top layer, respectively. The processing units 537 to 539 perform layering processing to interpolate the cropped region of the target block in latent space and generate a cropless image of the target block in latent space. Specifically, each of the processing units 537 to 539 performs convolution processing and upsizing processing using the target block information input from each of the combining units 540 to 542 as input to the corresponding layer. The processing units 537 and 538 each supply the resulting target block information to the combining units 541 and 542 in the layer one layer above themselves. The processing unit 539 converts the processing result into a cropless image of the target block in latent space and supplies it to the AE decoder 543.
[0286] The combining unit 540 combines the target block output from the AE encoder 534 with the attention information of the target block supplied from the raster conversion unit 536 , and supplies the resulting information of the target block to the processing unit 537 .
[0287] The combining unit 541 combines the target block output from the AE encoder 533 with the information on the target block supplied from the processing unit 537 , and supplies the resulting information on the target block to the processing unit 538 .
[0288] The combining unit 542 combines the target block output from the AE encoder 532 with the information on the target block supplied from the processing unit 538 , and supplies the resulting information on the target block to the processing unit 539 .
[0289] The AE decoder 543 (high-quality image decoding unit) converts the cropless image of the target block in the latent space supplied from the processing unit 539 into a cropless image of the target block in the pixel space and outputs it.
[0290] In this way, the image generation unit 511 divides the image in pixel space into processing unit blocks in raster scan order and generates a cropless image for each processing unit block. Therefore, cropless images can be generated faster than the image generation unit 117 in Fig. 14. Therefore, the image generation unit 511 is more advantageous than the image generation unit 117 when implemented in an embedded system.
[0291] <Explanation of Effects> FIG. 32 is a diagram illustrating the effects of the light crop interpolation process performed by the image generation unit 511 in FIG.
[0292] A in Fig. 32 shows, from top to bottom, the sizes of data processed in the topmost layer, second-topmost layer, and bottommost layer of the image generation unit 117 in Fig. 14. B in Fig. 32 shows, from top to bottom, the sizes of data processed in the topmost layer, second-topmost layer, and bottommost layer of the image generation unit 511 in Fig. 31.
[0293] The image generation unit 117 performs crop interpolation processing on an image-by-image basis (global window basis). Therefore, as shown in A of Fig. 32 , the image generation unit 117 processes the image from the electronic stabilization processing unit 113 in the latent space in the top layer. Therefore, in order to hold the input and processing results from the VQ-VAE encoder 114 in the top layer, a buffer with a size corresponding to the product of the first resolution, which is the resolution of the image, and the number of feature dimensions in the latent space is required.
[0294] In the image generation unit 117, an image of a second resolution lower than the first resolution in the latent space is processed in the second-highest layer. Therefore, in order to store the processing results in the second-highest layer, a buffer of a size corresponding to the product of the second resolution and the number of feature dimensions in the latent space is required.
[0295] In the image generation unit 117, an image with a third resolution lower than the second resolution is processed in the lowest layer. Therefore, in order to store the processing results in the lowest layer, a buffer with a size corresponding to the product of the third resolution and the number of feature dimensions of the latent space is required.
[0296] As described above, when calculating each layer, the image generation unit 117 needs a buffer with a size corresponding to the product of the resolution corresponding to that layer and the number of feature dimensions of the latent space. Therefore, although the crop interpolation process by the image generation unit 117 can generate a highly accurate cropless image, it is a process with high calculation costs that consumes a lot of memory such as a VRAM (Video Random Access Memory).
[0297] In contrast, the image generation unit 511 performs light crop interpolation processing on a processing unit block basis in raster scan order. Therefore, as shown in B of Fig. 32 , in the image generation unit 511, in the top layer, a processing unit block of the first resolution image is processed as a target block. Therefore, the size of the buffer required to hold the input from the AE encoder 532 and the processing results in the top layer is only a size corresponding to the product of the size of the processing unit block in the first resolution image and the number of feature dimensions in the latent space. In the example of Fig. 32 , the size of the processing unit block corresponds to the size of one line of the first resolution image, and the first target block in Fig. 32 is indicated by diagonal lines.
[0298] In the image generation unit 511, in the second-highest layer, a processing unit block in an image of a second resolution lower than the first resolution in the latent space is processed as a target block. Therefore, the size of the buffer required to hold the processing results in the second-highest layer only needs to be a size corresponding to the product of the size of the processing unit block in the image of the second resolution and the number of feature dimensions in the latent space.
[0299] In the image generation unit 511, in the lowest layer, a processing unit block in an image of a third resolution lower than the second resolution in the latent space is processed as a target block. Therefore, the size of the buffer required to hold the processing results in the lowest layer only needs to be a size corresponding to the product of the size of the processing unit block in the image of the third resolution and the number of feature dimensions in the latent space.
[0300] As described above, the image generation unit 511 can reduce the size of the buffer required for calculations at each layer compared to the image generation unit 117. Therefore, the calculation cost of the image generation unit 511 is smaller than the calculation cost of the image generation unit 117. Therefore, the light crop interpolation processing by the image generation unit 511 is advantageous when it is implemented as DSP processing in an embedded system.
[0301] In the first to sixth embodiments, the segment analysis classifies the image into segments of the ground, the sky, or segments other than the ground and the sky, but the image analysis model may have an extension function that enables classification into more detailed segments. This extension function enables effective processing using more detailed segment analysis information in depth analysis and motion analysis.
[0302] 7. Extended Functions of Image Analysis Model Example of Segment Analysis Information FIG. 33 is a diagram showing an example of segment analysis information in the first to sixth embodiments.
[0303] The left side of A to C in Figure 33 is an example of an image in pixel space whose resolution has been reduced by the reduction unit 21, and the right side is an example of segment analysis information generated from that image.
[0304] As shown in A to C of Fig. 33, the segment analysis information is information that represents the ground, sky, or other than the ground and sky as a segment of each pixel of the low-resolution image. In the example of Fig. 33, the segment of each pixel is represented by density in the segment analysis information, and the resolution of the low-resolution image and the segment analysis information is 64 x 64 pixels.
[0305] <Example of Segment Tree Structure> FIG. 34 is a diagram showing an example of a segment tree structure in an image analysis model having an extended function that enables classification into more detailed segments.
[0306] In the example of Figure 34, the image is first classified into sky, ground, or non-sky and non-ground segments. The ground segments are then further classified into indoor or outdoor segments, the non-sky and non-ground segments are classified into man-made structures, animals, or man-made structures and non-animals segments. The animal segments are further classified into human or non-human segments, and the human segments are further classified into face or non-face segments.
[0307] The segment tree structure is not limited to the example shown in FIG. 34, and any node can be added.
[0308] As described above, in an image analysis model with extended functions, the segments classified by segment analysis have a tree structure. Therefore, it is possible to perform only the extensions that enable classification of the necessary segments. As a result, it is possible to reduce the computational costs of image analysis processing and the development costs of the image analysis model.
[0309] 8. Computer The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs that make up the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, for example, that can execute various functions by installing various programs.
[0310] FIG. 35 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0311] In the computer, a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903 are interconnected by a bus 904.
[0312] An input / output interface 905 is also connected to the bus 904. An input unit 906, an output unit 907, a storage unit 908, a communication unit 909, and a drive 910 are connected to the input / output interface 905.
[0313] The input unit 906 includes a keyboard, a mouse, a microphone, etc. The output unit 907 includes a display, a speaker, etc. The storage unit 908 includes a hard disk, a non-volatile memory, etc. The communication unit 909 includes a network interface, etc. The drive 910 drives removable media 911 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0314] In a computer configured as described above, the CPU 901 performs the above-described series of processes by, for example, loading a program stored in the memory unit 908 into the RAM 903 via the input / output interface 905 and the bus 904 and executing it.
[0315] The program executed by the computer (CPU 901) can be provided by being recorded on a removable medium 911 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0316] In a computer, the program can be installed in the storage unit 908 via the input / output interface 905 by inserting the removable medium 911 into the drive 910. The program can also be received by the communication unit 909 via a wired or wireless transmission medium and installed in the storage unit 908. Alternatively, the program can be installed in the ROM 902 or the storage unit 908 in advance.
[0317] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0318] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0319] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.
[0320] For example, it is possible to adopt a form in which all or part of the above-described multiple embodiments are combined. For example, in the fifth embodiment, the imaging device 411 may generate an occlusion-less image as in the fourth embodiment, rather than a cropless image. As in the sixth embodiment, other processing performed by the image generation units 117 and 317, such as the occlusion interpolation processing in the fourth embodiment, may be reduced in weight.
[0321] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.
[0322] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0323] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0324] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0325] The present technology may have the following configurations: (1) An image processing device comprising: an image encoding unit that converts an image in pixel space into the image in a latent space; and an image analysis unit that performs multi-layer hierarchical encoding on the image in the latent space and performs multi-layer hierarchical decoding on the encoding result of a lowest layer of the multiple layers, thereby generating decoded results of each of the multiple layers as analysis information representing results of different types of analysis of the image in the latent space, wherein the image analysis unit is configured to decode each of the multiple layers on information combining the encoding result of that layer with the decoding result of the layer immediately below that layer. (2) The image processing device described in (1), wherein the type of analysis corresponding to the top layer of the multiple layers is segment analysis, the type of analysis corresponding to the layer immediately below the top layer of the multiple layers is depth analysis, and the type of analysis corresponding to the bottom layer is motion analysis. (3) The image processing device described in (2), wherein the segments in the segment analysis have a hexagonal tree structure. (4) The image processing device according to any one of (1) to (3), wherein the image analysis unit is configured to generate information combining the decoding result of a top layer of the multiple layers with the image in the latent space as a restored image by restoring the image in the latent space. (5) The image processing device according to (4), further comprising: an analysis information decoding unit that converts the analysis information in the latent space generated by the image analysis unit into analysis information in the pixel space, and an image decoding unit that converts the restored image in the latent space generated by the image analysis unit into the restored image in the pixel space. (6) The image processing device according to (1), wherein the image in the pixel space is generated by leaving a high-resolution image in the pixel space, having a resolution higher than that of the image in the pixel space, as is or by reducing the resolution.(7) The image processing device according to (6), further comprising: a high-resolution image encoding unit that converts the high-resolution image in the pixel space into the high-resolution image in the latent space; an image generation unit that improves the quality of the high-resolution image in the latent space using an image generation model that is a learning model to which the analysis information in the latent space generated by the image analysis unit is cross-attentionally connected, and generates a high-quality image in the latent space; and a high-quality image decoding unit that converts the high-quality image in the latent space generated by the image generation unit into the high-quality image in the pixel space. (8) The image processing device according to (7), further comprising: (9) The image processing device described in (7) above, further comprising: an electronic stabilization processing unit that performs electronic stabilization processing on the high-resolution image in the pixel space; wherein the high-resolution image encoding unit converts the high-resolution image after the electronic stabilization processing into the high-resolution image in the latent space; and the image generation unit uses the image generation model to generate an interpolated image in the latent space by interpolating at least one of a crop region and an occlusion region that has occurred in the high-resolution image by the electronic stabilization processing.(10) The image processing device according to (9), further comprising: a detection unit that detects at least one of acceleration and angular velocity of an imaging unit that captures the high-resolution image; a detection information analysis unit that generates detection information analysis information representing a result of an analysis of the detection information in the latent space from the detection information representing at least one of the acceleration and the angular velocity detected by the detection unit, using a detection information analysis model that is a learning model in which the analysis information in the latent space generated by the image analysis unit is cross-attentionally connected; and a detection information analysis information decoding unit that converts the detection information analysis information in the latent space generated by the detection information analysis unit into the detection information analysis information in the detection information space, wherein the electronic stabilization processing unit is configured to perform the electronic stabilization processing on the high-resolution image in the pixel space based on the detection information analysis information in the detection information space. (11) The image processing device according to (7) or (8), wherein the image generation unit is configured to use the image generation model to generate, as the high-quality image in the latent space, the high-resolution image that has been subjected to at least one of deblurring and noise removal in the latent space. (12) The image processing device according to any one of (7) to (11), further comprising: an analysis information decoding unit that converts the analysis information in the latent space generated by the image analysis unit into analysis information in the pixel space; and an output unit that outputs the high-quality image in the pixel space and the analysis information in the pixel space in association with each other.(13) The image processing apparatus further comprises: a pyramiding unit that generates a pyramid image consisting of images in the pixel space at multiple resolutions by downsampling the high-resolution image in the pixel space multiple times, divides each image in the pyramid image into processing unit blocks, and outputs them as target blocks in raster scan order; a pyramid image encoding unit that converts the target blocks output by the pyramiding unit into the target blocks in the latent space; an image generating unit that performs layering processing using the target blocks of each image in the pyramid image in the latent space as inputs of each layer to improve the quality of the target blocks in the high-resolution image in the latent space and generate a high-quality image of the target block in the latent space; and a high-quality image decoding unit that converts the high-quality image of the target block generated by the image generating unit into the high-quality image of the target block in the pixel space, The image processing device according to (6), wherein the processing of each layer other than the lowest layer in the layering processing is performed using as input information obtained by combining the input of that layer with the processing result of the layer one layer below that layer, and the processing of the lowest layer in the layering processing is performed using as input information obtained by combining the input of the lowest layer with the analytical information in the latent space generated by the image analysis unit.(14) The image processing device according to (13), further comprising a division unit that converts the analytical information in the two-dimensional latent space generated by the image analysis unit into analytical information in the one-dimensional latent space arranged in raster scan order, divides the analytical information into processing unit blocks, and outputs the converted analytical information as the analytical information of the target block in raster scan order, and wherein the processing of the lowest layer in the layering processing is performed using as input information obtained by combining the input of the lowest layer with the analytical information of the target block output from the division unit.(15) The image processing device according to (6), further comprising: a detection unit that detects at least one of acceleration and angular velocity of an imaging unit that captures the high-resolution image; a detection information analysis unit that generates detection information analysis information representing a result of an analysis of the detection information in the latent space from detection information representing at least one of the acceleration and the angular velocity detected by the detection unit, using a detection information analysis model that is a learning model in which the analysis information in the latent space generated by the image analysis unit is cross-attentionally connected; and a detection information analysis information decoding unit that converts the detection information analysis information in the latent space generated by the detection information analysis unit into the detection information analysis information in the detection information space. (16) The image processing device according to (15), further comprising: (17) The image processing device according to (15) or (16), wherein the detection information analysis information is composed of at least one of local motion information representing a movement of a subject in the high-resolution image, rotation / translation information consisting of rotation information representing a movement of the imaging unit in a rotational direction and translation information representing a movement of the imaging unit in a translational direction, the calibrated detection information, and attitude control information for controlling the attitude of the imaging unit. (18) An image processing method, wherein an image processing device includes: converting an image in a pixel space into the image in a latent space, performing multi-layer hierarchical encoding on the image in the latent space, and performing multi-layer hierarchical decoding on the encoding result of a lowest layer of the multiple layers, thereby generating decoded results of each of the multiple layers as analysis information representing results of different types of analysis of the image in the latent space, wherein the decoding of each of the multiple layers is performed on information combining the encoding result of that layer and the decoding result of the layer one layer below.(19) A program for causing a computer to execute a process including: converting an image in pixel space into the image in latent space; performing multi-layer hierarchical encoding on the image in the latent space; and performing multi-layer hierarchical decoding on the encoding result of the lowest layer of the multiple layers, thereby generating decoded results of each of the multiple layers as analysis information representing the results of different types of analysis of the image in the latent space, wherein the decoding of each of the multiple layers is performed on information combining the encoding result of that layer and the decoding result of the layer one layer below that layer. (20) An imaging system comprising: an imaging unit that captures an image in pixel space; an image encoding unit that converts the image in pixel space captured by the imaging unit into the image in latent space; and an image analysis unit that performs multi-layer hierarchical encoding on the image in the latent space and performs multi-layer hierarchical decoding on the encoding result of the lowest layer of the multiple layers, thereby generating decoded results of each of the multiple layers as analysis information representing the results of different types of analysis of the image in the latent space, wherein the image analysis unit is configured to decode each of the multiple layers using information that combines the encoding result of that layer with the decoding result of the layer one layer below it.
[0326] 10 Imaging system, 11 Imaging unit, 13 Image processing unit, 22 VAE encoder, 24 Image analysis unit, 25 VAE decoder, 26-28 Decoder, 110 Imaging system, 111 Imaging unit, 112 IMU sensor, 113 Electronic stabilization processing unit, 114 VQ-VAE encoder, 115 Image processing unit, 117 Image generation unit, 118 VQ-VAE decoder, 210 Imaging system, 212 IMU analysis unit, 213-217 Decoder, 310 Imaging system, 311 Rotational stabilization processing unit, 312 IMU analysis unit, 313 Image processing unit, 315 Translational stabilization processing unit, 317 Image generation unit, 411 Imaging device, 421 Image processing unit, 422 Image generation unit, 510 Imaging system, 531 Pyramid generation unit, 532 to 534 AE encoder, 536 Raster conversion unit, 537 to 539 Processing unit, 543 High-quality image decoding unit
Claims
an image encoder that converts an image in pixel space to said image in latent space; an image analysis unit that performs hierarchical encoding on the image in the latent space in multiple layers, and performs hierarchical decoding on the encoding result of the lowest layer of the multiple layers, thereby generating decoded results of each of the multiple layers as analysis information representing results of different types of analysis of the image in the latent space; Equipped with The image analysis unit decodes each of the plurality of layers using information obtained by combining the encoding result of that layer with the decoding result of the layer immediately below that layer. It was configured as Image processing device. the type of analysis corresponding to the top layer of the plurality of layers is segment analysis; the type of analysis corresponding to the layer one layer below the top layer among the plurality of layers is a depth analysis; The type of analysis corresponding to the lowest layer is motion analysis. It was configured as The image processing device according to claim 1 . The segments in the segment analysis have a branching structure. It was configured as The image processing device according to claim 2 . The image analysis unit generates information combining the decoding result of the top layer of the plurality of layers and the image in the latent space as a restored image obtained by restoring the image in the latent space. It was configured as The image processing device according to claim 1 . an analysis information decoding unit that converts the analysis information in the latent space generated by the image analysis unit into analysis information in the pixel space; an image decoding unit that converts the restored image in the latent space generated by the image analysis unit into the restored image in the pixel space; Further equipped The image processing device according to claim 4 . The image in the pixel space is generated by retaining or reducing a high-resolution image in the pixel space, the high-resolution image having a higher resolution than the resolution of the image in the pixel space. It was configured as The image processing device according to claim 1 . a high-resolution image encoding unit that converts the high-resolution image in the pixel space into the high-resolution image in the latent space; an image generation unit that improves the quality of the high-resolution image in the latent space using an image generation model that is a learning model in which the analysis information in the latent space generated by the image analysis unit is cross-attention connected, and generates a high-quality image in the latent space; a high-quality image decoding unit that converts the high-quality image in the latent space generated by the image generating unit into the high-quality image in the pixel space; Further equipped The image processing device according to claim 6 . the image analysis unit generates information combining the decoding result of the top layer of the plurality of layers with the image in the latent space as a restored image obtained by restoring the image in the latent space; The image generation model is also connected to the reconstructed image in the latent space by cross-attention. It was configured as The image processing device according to claim 7 . an electronic stabilization processing unit that performs electronic stabilization processing on the high-resolution image in the pixel space; Furthermore, the high-resolution image encoding unit converts the high-resolution image after the electronic stabilization processing into the high-resolution image in the latent space; The image generation unit uses the image generation model to generate an interpolated image in the latent space by interpolating at least one of a crop region and an occlusion region that has occurred in the high-resolution image by the electronic stabilization process. It was configured as The image processing device according to claim 7 . a detection unit that detects at least one of acceleration and angular velocity of an imaging unit that captures the high-resolution image; a detection information analysis unit that generates detection information analysis information representing a result of an analysis of the detection information in the latent space from detection information representing at least one of the acceleration and the angular velocity detected by the detection unit, using a detection information analysis model that is a learning model in which the analysis information in the latent space generated by the image analysis unit is cross-attention connected; a detection information analysis information decoding unit that converts the detection information analysis information in the latent space generated by the detection information analysis unit into detection information analysis information in a detection information space; Furthermore, The electronic stabilization processing unit performs the electronic stabilization processing on the high-resolution image in the pixel space based on the detection information analysis information in the detection information space. It was configured as The image processing device according to claim 9 . The image generation unit generates, using the image generation model, the high-resolution image in the latent space that has been subjected to at least one of deblur correction and noise removal, as the high-quality image in the latent space. It was configured as The image processing device according to claim 7 . an analysis information decoding unit that converts the analysis information in the latent space generated by the image analysis unit into analysis information in the pixel space; an output unit that outputs the high-quality image in the pixel space and the analysis information in the pixel space in association with each other; Further equipped The image processing device according to claim 7 . a pyramiding unit that generates a pyramid image consisting of images in the pixel space at a plurality of resolutions by downsampling the high-resolution image in the pixel space a plurality of times, divides each image in the pyramid image into processing unit blocks, and outputs the divided images as target blocks in a raster scan order; a pyramid image encoding unit that converts the target block output by the pyramiding unit into the target block in the latent space; an image generation unit that performs a hierarchical process using the target block of each image of the pyramid image in the latent space as an input of each layer to improve the quality of the target block of the high-resolution image in the latent space and generate a high-quality image of the target block in the latent space; a high-quality image decoding unit that converts the high-quality image of the target block generated by the image generating unit into the high-quality image of the target block in the pixel space; Furthermore, The processing of each layer other than the lowest layer in the layering process is performed using as input information that combines the input of that layer with the processing result of the layer one layer below that layer, and the processing of the lowest layer in the layering process is performed using as input information that combines the input of that lowest layer with the analysis information in the latent space generated by the image analysis unit. The image processing device according to claim 6 . a division unit that converts the analysis information in the two-dimensional latent space generated by the image analysis unit into analysis information in the one-dimensional latent space arranged in raster scan order, divides the analysis information into processing unit blocks, and outputs the analysis information of the target block in raster scan order. Furthermore, The processing of the lowest layer in the layering processing is performed using as input information that combines the input of the lowest layer and the analysis information of the target block output from the dividing unit. It was configured as The image processing device according to claim 13. a detection unit that detects at least one of acceleration and angular velocity of an imaging unit that captures the high-resolution image; a detection information analysis unit that generates detection information analysis information representing a result of an analysis of the detection information in the latent space from detection information representing at least one of the acceleration and the angular velocity detected by the detection unit, using a detection information analysis model that is a learning model in which the analysis information in the latent space generated by the image analysis unit is cross-attention connected; a detection information analysis information decoding unit that converts the detection information analysis information in the latent space generated by the detection information analysis unit into detection information analysis information in a detection information space; Further equipped The image processing device according to claim 6 . the image analysis unit generates information combining the decoding result of the top layer of the plurality of layers with the image in the latent space as a restored image obtained by restoring the image in the latent space; The detection information analysis model is also connected to the restored image in the latent space by cross-attention. It was configured as The image processing device according to claim 15. The detection information analysis information is composed of at least one of local motion information representing the movement of the subject in the high-resolution image, rotation / translation information consisting of rotation information representing the movement of the imaging unit in the rotation direction and translation information representing the movement in the translation direction, the calibrated detection information, and attitude control information for controlling the attitude of the imaging unit. The image processing device according to claim 15. The image processing device Transforming an image in pixel space to said image in latent space; performing a multi-layer hierarchical encoding on the image in the latent space, and performing a multi-layer hierarchical decoding on the encoding result of the lowest layer of the multiple layers, thereby generating the decoding results of each of the multiple layers as analysis information representing the results of different types of analysis of the image in the latent space; Including, The decoding of each layer of the plurality of layers is performed on information obtained by combining the encoding result of that layer with the decoding result of the layer immediately below that layer. Image processing methods. Transforming an image in pixel space to said image in latent space; performing a multi-layer hierarchical encoding on the image in the latent space, and performing a multi-layer hierarchical decoding on the encoding result of the lowest layer of the multiple layers, thereby generating the decoding results of each of the multiple layers as analysis information representing the results of different types of analysis of the image in the latent space; Including, The decoding of each layer of the plurality of layers is performed on information obtained by combining the encoding result of that layer with the decoding result of the layer immediately below that layer. A program that causes a computer to execute a process. an imaging unit that captures an image in pixel space; an image encoding unit that converts the image in the pixel space captured by the imaging unit into an image in a latent space; an image analysis unit that performs hierarchical encoding on the image in the latent space in multiple layers, and performs hierarchical decoding on the encoding result of the lowest layer of the multiple layers, thereby generating decoded results of each of the multiple layers as analysis information representing results of different types of analysis of the image in the latent space; Equipped with The image analysis unit decodes each of the plurality of layers using information obtained by combining the encoding result of that layer with the decoding result of the layer immediately below that layer. It was configured as Imaging system.
Citation Information
Patent Citations
Information processing device and information processing method
JP2021058563A
Improved medical scanning protocols for in-scanner patient data acquisition and analysis
JP2022539063A
Determination of cardiac functional indices
WO2023199088A1