Learning device, video image generation device, trained model generation method, video image generation method, and program

The learning device and method address the challenge of uniformly expanding the dynamic range and maintaining temporal consistency in HDR image generation by using a machine learning model with forward and backward processes, achieving high-quality HDR images from SDR inputs.

JP7780465B2Active Publication Date: 2025-12-04SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022581151
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-02-15
Publication Date
2025-12-04
Estimated Expiration
2041-02-15

AI Technical Summary

Technical Problem

Existing technologies struggle to uniformly and accurately expand the dynamic range of standard dynamic range (SDR) images to high dynamic range (HDR) images, particularly in moving images, leading to issues with pixel saturation and temporal consistency.

Method used

A learning device and method that utilizes a machine learning model, incorporating a deep neural network with a recurrent neural network, to generate HDR images by associating luminance values in SDR and HDR through correspondence data, employing forward and backward generation processes to ensure accurate dynamic range expansion and temporal consistency.

Benefits of technology

The solution effectively expands the dynamic range of SDR images to HDR while maintaining temporal consistency, ensuring accurate pixel luminance values and reducing the need for separate models for forward and backward processes, thus enhancing the quality of generated HDR moving images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007780465000001
    Figure 0007780465000001
  • Figure 0007780465000002
    Figure 0007780465000002
  • Figure 0007780465000003
    Figure 0007780465000003
Patent Text Reader

Abstract

Provided are a learning device, a learned model generating method and a program wherein an increase in dynamic range can be appropriately achieved for a variety of images in a unified manner. A training data generating unit (64), by referring to correspondence data in which brightness values in a second dynamic range are associated with brightness values in a first dynamic range, generates, on the basis of a second type of image, a first type of image associated with the second type of image. A learning unit (68) executes learning of a machine learning model (30) by using the first type of image and the second type of image associated with the first type of image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a learning device, a moving image generating device, a trained model generating method, a moving image generating method, and a program. [Background technology]

[0002] In recent years, high dynamic range (HDR) moving images have been generated by expanding the dynamic range of standard dynamic range (SDR) moving images, which are legacy video assets.

[0003] As a related technology, Non-Patent Document 1 describes a technology that uses a trained convolutional neural network (CNN) to generate an HDR image by expanding the dynamic range of an SDR image. Note that Non-Patent Document 1 refers to SDR as LDR (low dynamic range).

[0004] In the technology described in Non-Patent Document 1, an SDR image in which the luminance values ​​of 5 percent of pixels are saturated is generated from an existing HDR image. Then, CNN learning is performed so that the original HDR image can be reproduced from the generated SDR image. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Gabriel Eilertsen and 3 others, "HDR image reconstruction from a single exposure using deep CNNs", [online], October 20, 2017, ACM Transactions on Graphics, Vol.36, No.6, Article 178, [searched on February 1, 2021], Internet <URL: https: / / arxiv.org / abs / 1710.07480> Summary of the Invention [Problem to be solved by the invention]

[0006] The technology described in Non-Patent Document 1 generates an SDR image in which 5% of the pixels have a uniform saturated brightness value, regardless of the scene of the HDR image, as described above. However, the percentage of pixels with saturated brightness values ​​varies in actual SDR images. Therefore, a CNN trained on such SDR images cannot uniformly and accurately expand the dynamic range of various actual SDR images.

[0007] Furthermore, since the technology described in Non-Patent Document 1 is not intended for moving images, using the technology described in Non-Patent Document 1 may result in a decrease in the temporal consistency of the generated HDR moving images.

[0008] Therefore, if a recurrent neural network (RNN) is used to generate each frame image based on the frame image of the previous frame that has already been generated, the temporal consistency of the generated HDR video can be improved. However, in this case, the dynamic range of the first frame image cannot be expanded accurately because there is no frame image of the previous frame.

[0009] The present invention has been made in consideration of the above-mentioned situation, and one of its purposes is to provide a learning device, a method for generating a trained model, and a program that can uniformly and accurately expand the dynamic range for various images.

[0010] Another object of the present invention is to provide a moving image generating device, a moving image generating method, and a program that can accurately expand the dynamic range of the first frame while ensuring the temporal consistency of the generated moving image. [Means for solving the problem]

[0011] In order to solve the above problem, the learning device of the present invention is a learning device that executes learning of a machine learning model that outputs, in response to input of a first type of image that is an image of a first dynamic range, a second type of image that is an image of a second dynamic range obtained by widening the dynamic range of the first type of image, and includes: an image generation unit that generates, based on the second type of image, a first type of image that is associated with the second type of image by referring to correspondence data in which luminance values ​​in the second dynamic range are associated with luminance values ​​in the first dynamic range; and a learning unit that executes learning of the machine learning model using the first type of image and the second type of image that is associated with the first type of image.

[0012] In one aspect of the present invention, the correspondence data is data in which luminance values ​​equal to or greater than a predetermined value in the second dynamic range are associated with saturation values ​​in the first dynamic range.

[0013] In one aspect of the present invention, the correspondence data is a lookup table.

[0014] The correspondence data may be a one-dimensional lookup table.

[0015] In one aspect of the present invention, the first dynamic range is a standard dynamic range (SDR), and the second dynamic range is a high dynamic range (HDR).

[0016] Furthermore, a moving image generation device according to the present invention is a moving image generation device that uses a trained model to generate, based on a first type moving image that is a moving image with a first dynamic range, a second type moving image that is a moving image with a second dynamic range obtained by widening the dynamic range of the first type moving image, wherein the trained model is a trained model that generates, based on a first type frame image that is a frame image included in the first type moving image and a second type frame image that is a frame image with the second dynamic range that is adjacent to the first type frame image, the second type frame image that is associated with the first type frame image, and the moving image generation device uses the trained model to generate a direct correlation between the first type frame image and the first type frame image generated by the trained model. The system includes a forward generation process execution unit that executes forward generation process to generate the second-type frame image associated with the first-type frame image based on the second-type frame image of the previous frame; a backward generation process execution unit that executes backward generation process to generate the second-type frame image associated with the first-type frame image, for at least the first frame, based on the first-type frame image and the second-type frame image of the frame immediately following the first-type frame image generated by the trained model, using the trained model; and a moving image generation unit that generates the second-type moving image based on the second-type frame image generated by the forward generation process and the second-type frame image generated by the backward generation process.

[0017] In one aspect of the present invention, the trained model is a trained model that generates the second type of frame image associated with a first type of frame image based on the first type of frame image, the first type of frame image of a frame adjacent to the first type of frame image, and the second type of frame image of the adjacent frame, and the forward generation process execution unit generates the second type of frame image associated with the first type of frame image based on the first type of frame image, the first type of frame image of a frame immediately preceding the first type of frame image, and the second type of frame image of the immediately preceding frame generated by the trained model, and the backward generation process execution unit generates the second type of frame image associated with the first type of frame image based on the first type of frame image, the first type of frame image of a frame immediately following the first type of frame image, and the second type of frame image of the immediately following frame generated by the trained model.

[0018] In addition, in one aspect of the present invention, the backward generation process execution unit generates the second type of frame image corresponding to the first type of frame image based on the first type of frame image and the second type of frame image that is the frame immediately following the first type of frame image generated in the forward generation process.

[0019] Alternatively, the backward generation processing execution unit generates the second type frame image that corresponds to the first type frame image based on the first type frame image and the second type frame image that is the frame immediately following the first type frame image generated by the backward generation processing.

[0020] In addition, in one aspect of the present invention, the video generation unit generates the second type of video including the second type of frame image of the first frame generated in the backward generation process and the second type of frame images of the second and subsequent frames generated in the forward generation process.

[0021] Alternatively, for at least one frame, the video generation unit generates a frame image of the frame included in the second type video based on the second type frame image of the frame generated in the forward generation process and the second type frame image of the frame generated in the backward generation process.

[0022] In this aspect, the moving image generation unit may generate the second type moving image for at least one frame in which the weighted average of the luminance value of a pixel included in the second type frame image of the frame generated by the forward generation process and the luminance value of the pixel included in the second type frame image of the frame generated by the backward generation process is set to the luminance value of the pixel included in the frame image of the frame.

[0023] In one aspect of the present invention, the first dynamic range is a standard dynamic range (SDR), and the second dynamic range is a high dynamic range (HDR).

[0024] Furthermore, a method for generating a trained model according to the present invention is a method for generating a trained model that executes training of a machine learning model that outputs, in response to input of a first type of image that is an image of a first dynamic range, a second type of image that is an image of a second dynamic range obtained by widening the dynamic range of the first type of image, and includes the steps of: generating, based on the second type of image, a first type of image that is associated with the second type of image by referring to correspondence data in which luminance values ​​in the second dynamic range are associated with luminance values ​​in the first dynamic range; and executing training of the machine learning model using the first type of image and the second type of image that is associated with the first type of image.

[0025] Furthermore, a moving image generation method according to the present invention is a moving image generation method that uses a trained model to generate, based on a first type moving image that is a moving image with a first dynamic range, a second type moving image that is a moving image with a second dynamic range obtained by widening the dynamic range of the first type moving image, wherein the trained model is a trained model that generates, based on a first type frame image that is a frame image included in the first type moving image and a second type frame image that is a frame image of the second dynamic range that is adjacent to the first type frame image, the second type frame image that is associated with the first type frame image, and the moving image generation method uses the trained model to generate a second type frame image that is associated with the first type frame image and the first type frame image generated by the trained model. the step of executing a forward generation process to generate the second-type frame image associated with the first-type frame image based on the second-type frame image of the frame immediately preceding the first-type frame image; the step of executing a backward generation process to generate the second-type frame image associated with the first-type frame image based on the first-type frame image and the second-type frame image of the frame immediately following the first-type frame image generated by the trained model, for at least the first frame, using the trained model; and the step of generating the second-type moving image based on the second-type frame image generated by the forward generation process and the second-type frame image generated by the backward generation process.

[0026] Furthermore, the program according to the present invention causes a computer that executes training of a machine learning model that outputs, in response to input of a first type image that is an image of a first dynamic range, a second type image that is an image of a second dynamic range obtained by widening the dynamic range of the first type image, to execute the following steps: generating, based on the second type image, a first type image that is associated with the second type image by referring to correspondence data in which luminance values ​​in the second dynamic range are associated with the first dynamic range; and executing training of the machine learning model using the first type image and the second type image that is associated with the first type image.

[0027] Another program according to the present invention is a program that causes a computer to use a trained model to generate, based on a first type of video, which is a video having a first dynamic range, a second type of video that is a video having a second dynamic range obtained by widening the dynamic range of the first type of video, wherein the trained model is a trained model that generates, based on a first type of frame image that is a frame image included in the first type of video and a second type of frame image that is a frame image of the second dynamic range that is adjacent to the first type of frame image, the second type of frame image that is associated with the first type of frame image, and the program causes the computer to use the trained model to generate, based on the first type of frame image and the second type of frame image that is adjacent to the first type of frame image, the second type of frame image that is associated with the first type of frame image, the trained model is used to execute a forward generation process to generate the second-type frame image associated with the first-type frame image based on the second-type frame image of the frame immediately preceding the first-type frame image generated by the trained model; the trained model is used to execute a backward generation process to generate the second-type frame image associated with the first-type frame image for at least the first frame based on the first-type frame image and the second-type frame image of the frame immediately following the first-type frame image generated by the trained model; and the trained model is used to generate the second-type moving image based on the second-type frame image generated by the forward generation process and the second-type frame image generated by the backward generation process. [Brief explanation of the drawings]

[0028] [Figure 1] 1 is a diagram illustrating an example of a configuration of an image processing apparatus according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram illustrating an example of a data structure of training data. [Figure 3] FIG. 1 is a diagram illustrating an example of correspondence between luminance values ​​in HDR and luminance values ​​in SDR. [Figure 4] FIG. 1 is a diagram illustrating an example of learning of a machine learning model. [Figure 5] FIG. 1 is a diagram illustrating an example of learning of a machine learning model. [Figure 6] FIG. 10 is a diagram illustrating an example of generating a target HDR moving image. [Figure 7] FIG. 10 is a diagram illustrating an example of a forward generation process. [Figure 8] FIG. 10 is a diagram illustrating an example of a backward generation process. [Figure 9] FIG. 2 is a functional block diagram showing an example of functions implemented in an image processing device according to an embodiment of the present invention. [Figure 10] FIG. 10 is a diagram illustrating an example of generating a target HDR moving image. [Figure 11] FIG. 10 is a diagram illustrating an example of a backward generation process. [Figure 12] FIG. 2 is a flowchart showing an example of a flow of processing performed in an image processing apparatus according to an embodiment of the present invention. [Figure 13] FIG. 2 is a flowchart showing an example of a flow of processing performed in an image processing apparatus according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0029] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.

[0030] 1 is a configuration diagram of an image processing device 10 according to this embodiment. The image processing device 10 according to this embodiment is, for example, a computer such as a server computer, a personal computer, or a game console. As shown in FIG. 1, the image processing device 10 according to this embodiment includes, for example, a processor 12, a storage unit 14, an operation unit 16, and a display unit 18.

[0031] The processor 12 is a program-controlled device such as a CPU that operates according to a program installed in the image processing device 10, for example.

[0032] The storage unit 14 is a storage element such as a ROM or RAM, a hard disk drive, a solid state drive, etc. The storage unit 14 stores programs executed by the processor 12, etc.

[0033] The operation unit 16 is a user interface such as a keyboard, a mouse, or a game console controller, and receives operation inputs from the user and outputs signals indicating the contents of the inputs to the processor 12.

[0034] The display unit 18 is a display device such as a liquid crystal display, and displays various images according to instructions from the processor 12.

[0035] The image processing device 10 may include a communication interface such as a network board, an optical disc drive that reads optical discs such as DVD-ROMs and Blu-ray (registered trademark) discs, a USB (Universal Serial Bus) port, and the like.

[0036] A trained machine learning model is implemented in the image processing device 10 according to this embodiment. The image processing device 10 according to this embodiment uses the machine learning model to generate high dynamic range (HDR) moving images, which are obtained by expanding the dynamic range of standard dynamic range (SDR) moving images, based on the SDR moving images.

[0037] Hereinafter, SDR moving images will be referred to as SDR moving images, and HDR moving images will be referred to as HDR moving images. Furthermore, frame images included in SDR moving images will be referred to as SDR frame images, and frame images included in HDR moving images will be referred to as HDR frame images. Furthermore, the color space of SDR moving images according to this embodiment is, for example, Rec709 / r2.4, and the color space of HDR moving images according to this embodiment is, for example, Rec2020 / PQ.

[0038] An example of learning of the machine learning model implemented in the image processing device 10 will be described below.

[0039] In training the machine learning model according to this embodiment, first, training data 20, the data structure of which is exemplified in FIG. 2, is generated based on a given HDR video image.

[0040] 2, training data 20 according to this embodiment includes training input SDR videos 22 and teacher HDR videos 24. The training input SDR videos 22 are input to a machine learning model during training of the machine learning model. The teacher HDR videos 24 are used as teacher data during training of the machine learning model.

[0041] Hereinafter, the SDR frame images included in the training input SDR video 22 will be expressed as SDR frame image a(a(0), a(1), a(2), a(3), ...). The numbers in parentheses correspond to the frame numbers; for example, the SDR frame image a of the kth frame will be expressed as SDR frame image a(k-1).

[0042] Furthermore, the HDR frame images included in the teacher HDR video 24 will be expressed as HDR frame images b(b(0), b(1), b(2), b(3), ...). The numbers in parentheses correspond to the frame numbers; for example, the HDR frame image b of the kth frame will be expressed as HDR frame image b(k-1).

[0043] In the present embodiment, for example, by referencing correspondence data in which luminance values ​​in HDR and luminance values ​​in SDR are associated with a given teacher HDR moving image 24, a learning input SDR moving image 22 associated with the given teacher HDR moving image 24 is generated based on the given teacher HDR moving image 24. For example, for each pixel included in the HDR frame image b of the given teacher HDR moving image 24, the luminance value of the pixel is converted to a luminance value in SDR, and color conversion is performed using a 3 × 3 matrix, thereby generating a learning input SDR moving image 22 associated with the given teacher HDR moving image 24.

[0044] Here, for example, the teacher HDR moving image 24 may be a moving image showing one scene. For example, an HDR moving image of a movie or the like may be divided into scenes based on an EDL (Edit Decision List), and each of the resulting moving images may be used as the teacher HDR moving image 24.

[0045] The correspondence data according to the present embodiment may be a lookup table such as a one-dimensional lookup table (1D-LUT), etc. The correspondence data may indicate the correspondence of luminance values ​​taking gamma correction into consideration.

[0046] Fig. 3 is a diagram showing an example of the correspondence between HDR luminance values ​​P and SDR luminance values ​​Q, as shown in the correspondence data according to this embodiment. In the example of Fig. 2, luminance values ​​equal to or greater than a predetermined value P1 in HDR are associated with a saturation value Q1 for luminance values ​​in SDR. In this way, this embodiment generates an SDR moving image in which the luminance values ​​of all pixels in the frame images of the HDR moving image that have luminance values ​​equal to or greater than P1 are converted to the saturation value Q1.

[0047] In this embodiment, by referring to the above-described correspondence data for each of the plurality of teacher HDR videos 24, a learning input SDR video 22 associated with the teacher HDR video 24 is generated. Note that the number of frames of the plurality of teacher HDR videos 24 may or may not be the same.

[0048] Then, for each of the plurality of teacher HDR videos 24, a plurality of sets of training data 20 including the teacher HDR video 24 and the learning input SDR video 22 associated with the teacher HDR video 24 are generated.

[0049] Furthermore, in this embodiment, in order to increase the number of training data 20, for example, a video may be generated by playing back a given teacher HDR video 24 in reverse (i.e., a video in which the frame order is reversed from that of the teacher HDR video 24). The video generated in this manner may also be used as the teacher HDR video 24.

[0050] Furthermore, for example, a moving image may be generated by connecting a moving image obtained by forward playing a given teacher HDR moving image 24 with a moving image obtained by reverse playing the given teacher HDR moving image 24. In this case, a moving image having twice the number of frames as the given teacher HDR moving image 24 is generated. The moving image generated in this manner may also be used as the teacher HDR moving image 24.

[0051] In this embodiment, the plurality of training data 20 thus generated is used to perform learning of the machine learning model 30 illustrated in FIG.

[0052] The machine learning model 30 according to this embodiment is a deep neural network (DNN) incorporating the mechanism of a recurrent neural network (RNN), and input to the machine learning model 30 is performed sequentially multiple times. Then, the output from the machine learning model 30 is included as part of the next input to the machine learning model 30.

[0053] As shown in FIG. 4, the machine learning model 30 according to this embodiment includes a first concatenation block 32, a feature extraction block 34, a resizing block 36, a second concatenation block 38, and an image generation block 40.

[0054] In this embodiment, for each SDR frame image a included in the training input SDR video 22, the machine learning model 30 generates an HDR frame image that is associated with the SDR frame image a.

[0055] Hereinafter, as shown in FIGS. 4 and 5, HDR frame images output from the machine learning model 30 during learning of the machine learning model 30 will be expressed as HDR frame images c(c(0), c(1), c(2), c(3), ...). Note that the numbers in parentheses correspond to the frame numbers; for example, the HDR frame image c of the kth frame will be expressed as HDR frame image c(k-1). Furthermore, as shown in FIG. 5, an HDR video including HDR frame image c generated in this manner will be referred to as a reference HDR video 50.

[0056] For example, an HDR frame image c(t) included in the reference HDR video 50 is generated based on an SDR frame image a(t) included in the training input SDR video 22. Note that in this embodiment, for example, the number of vertical and horizontal pixels of the SDR frame image a is the same as the number of vertical and horizontal pixels of the HDR frame image c associated with the SDR frame image.

[0057] In this embodiment, for example, the SDR frame images a are input to the machine learning model 30 sequentially in the order of frame numbers, starting from the SDR frame image a(0) of the first frame.

[0058] In this embodiment, when generating an HDR frame image c(t), not only the SDR frame image a(t) but also the SDR frame image a(t-1) of the immediately preceding frame is input to the machine learning model 30. In addition, the HDR frame image c(t-1), which is the output of the immediately preceding frame, is also input to the machine learning model 30.

[0059] 4, for example, in this embodiment, the first connection block 32 connects the SDR frame image a(t-1) and the HDR frame image c(t-1) to generate first intermediate data 42. Then, the first intermediate data 42 is input to the feature extraction block 34.

[0060] The feature extraction block 34 corresponds to, for example, a convolution layer or a pooling layer of a convolutional neural network (CNN), and outputs a feature map 44 in response to input of first intermediate data 42. The feature map 44 is data corresponding to, for example, an image (map) that is the output of the convolution layer or the pooling layer.

[0061] Then, the resizing block 36 enlarges the feature map 44 output from the feature extraction block 34 to the size (number of vertical and horizontal pixels) of the SDR frame image, thereby generating an enlarged feature map 46.

[0062] Then, the second connection block 38 connects the SDR frame image a(t) and the enlarged feature map 46 to generate second intermediate data 48. The second intermediate data 48 is then input to the image generation block 40.

[0063] The image generation block 40 is, for example, a convolutional neural network (CNN), and outputs an HDR frame image c(t) in response to input of the second intermediate data 48. In this way, the HDR frame image c(t) is generated. Then, as described above, this HDR frame image c(t) is included in the input to the machine learning model 30 when generating an HDR frame image c(t+1) that corresponds to the SDR frame image a(t+1).

[0064] In this embodiment, when generating the HDR frame image c(0), in addition to the SDR frame image a(0), a dummy SDR frame image a(-1) and a dummy HDR frame image c(-1) may be input to the machine learning model 30. Here, an image in which all pixels have the same brightness value may be used as the dummy image. For example, an image in which all pixels are white pixels or an image in which all pixels are black pixels may be used as the dummy image.

[0065] When the above process is performed for all SDR frame images a to generate the reference HDR video 50, the error (comparison result) between the reference HDR video 50 and the teacher HDR video 24 is identified, as shown in Fig. 5. Then, supervised learning is performed to update the parameter values ​​of the machine learning model 30 by backpropagation so as to minimize the value of the loss function associated with the identified error. In this embodiment, for example, supervised learning is performed using a known loss function that aims for time-series stability.

[0066] In this embodiment, for example, the above-described process is performed on a plurality of training data 20 related to video images of various scenes, thereby training the machine learning model 30. In this manner, a trained machine learning model 30 is generated.

[0067] Then, using the trained machine learning model 30 that has undergone the above-described training, for example, HDR moving images are generated by expanding the dynamic range of SDR moving images that are past video assets.

[0068] An example of generating HDR moving images using the trained machine learning model 30 will be described below.

[0069] 6, the SDR video that is the target for generating an HDR video will be referred to as a target SDR video 52. Furthermore, the HDR video that is generated by converting the target SDR video 52 into HDR video using the trained machine learning model 30 will be referred to as a target HDR video 54.

[0070] Furthermore, the SDR frame images included in the target SDR video 52 will be expressed as SDR frame images x(x(0), x(1), x(2), x(3), ...). The numbers in parentheses correspond to the frame numbers; for example, the SDR frame image x of the kth frame will be expressed as SDR frame image x(k-1).

[0071] In generating the target HDR video 54 according to this embodiment, two types of processing are executed: forward generation processing and backward generation processing.

[0072] Hereinafter, the HDR frame image generated by the forward generation process will be referred to as forward HDR frame image g1 (g1(0), g1(1), g1(2), g1(3), ...). Note that the number in parentheses corresponds to the frame number; for example, the forward HDR frame image g1 of the kth frame will be referred to as forward HDR frame image g1(k-1).

[0073] Furthermore, the HDR frame image generated by the reverse generation process will be referred to as a reverse HDR frame image g2. Note that in the example of the reverse generation process described below, a reverse HDR frame image g2 is generated only for the first frame. This reverse HDR frame image g2 will be referred to as a reverse HDR frame image g2(0).

[0074] Then, a target HDR moving image 54 is generated based on the forward HDR frame image g1 and the reverse HDR frame image g2.

[0075] Hereinafter, the HDR frame images included in the target HDR video 54 will be expressed as HDR frame images g (g(0), g(1), g(2), g(3), ...). The numbers in parentheses correspond to the frame numbers; for example, the HDR frame image g of the kth frame will be expressed as HDR frame image g(k-1).

[0076] In the forward generation process according to this embodiment, the SDR frame images x are input to the machine learning model 30 sequentially in the order of frame numbers, starting with the SDR frame image x(0) of the first frame.

[0077] 7, in the forward generation process, when a forward HDR frame image g1(t) is generated, not only the SDR frame image x(t) but also the SDR frame image x(t-1) of the immediately preceding frame is input to the machine learning model 30. In addition, the forward HDR frame image g1(t-1), which is the output of the immediately preceding frame, is also input to the machine learning model 30.

[0078] Then, the first connection block 32 connects the SDR frame image x(t-1) and the forward HDR frame image g1(t-1) to generate first intermediate data 42. The first intermediate data 42 is then input to the feature extraction block 34.

[0079] Then, the feature extraction block 34 and the resizing block 36 perform the same processing as in the training of the machine learning model 30 described above, thereby generating an enlarged feature map 46.

[0080] Then, the second connection block 38 connects the SDR frame image x(t) and the enlarged feature map 46 to generate second intermediate data 48. The second intermediate data 48 is then input to the image generation block 40.

[0081] Then, the image generation block 40 outputs the forward HDR frame image g1(t) in response to the input of the second intermediate data 48. In this way, the forward HDR frame image g1(t) is generated.

[0082] Note that in this embodiment, when generating the forward HDR frame image g1(0), in addition to the SDR frame image x(0), an SDR frame image x(-1) that is a dummy image and a forward HDR frame image g1(-1) that is also a dummy image may be input to the machine learning model 30. Here, an image in which all pixels have the same luminance value may be used as the dummy image. For example, an image in which all pixels are white pixels or an image in which all pixels are black pixels may be used as the dummy image.

[0083] In this embodiment, the above process is performed on all SDR frame images x in order, starting from the first frame, to generate multiple forward HDR frame images g1 (g1(0), g1(1), g1(2), g1(3), ...).

[0084] As described above, in the forward generation process, a dummy image is used when generating the forward HDR frame image g1(0) of the first frame. Therefore, the forward HDR frame image g1(0) of the first frame is estimated with lower accuracy by the trained machine learning model 30 than the forward HDR frame images g1 of the other frames.

[0085] Taking this into consideration, in this embodiment, the backward generation process is performed on the first frame as shown in FIG.

[0086] In the backward generation process, as shown in FIG. 8, not only the SDR frame image x(0) of the first frame included in the target SDR video 52, but also the SDR frame image x(1) of the second frame are input to the machine learning model 30. In addition, the forward HDR frame image g1(1) of the second frame is also input to the machine learning model 30. In this way, the forward HDR frame image g1(1) output from the machine learning model 30 by the forward generation process may be used as input to the machine learning model 30 in the backward generation process.

[0087] Then, the first connection block 32 connects the SDR frame image x(1) and the forward HDR frame image g1(1) to generate first intermediate data 42. The first intermediate data 42 is then input to the feature extraction block 34.

[0088] Then, the feature extraction block 34 and the resizing block 36 perform the same processing as in the training of the machine learning model 30 described above, thereby generating an enlarged feature map 46.

[0089] Then, the second connection block 38 connects the SDR frame image x(0) and the enlarged feature map 46 to generate second intermediate data 48. The second intermediate data 48 is then input to the image generation block 40.

[0090] Then, the image generation block 40 outputs the reverse HDR frame image g2(0) in response to the input of the second intermediate data 48. In this way, the reverse HDR frame image g2(0) for the first frame is generated.

[0091] Then, as described above, the target HDR video 54 is generated based on the forward HDR frame image g1 generated in the forward generation process and the backward HDR frame image g2 generated in the backward generation process.

[0092] For example, a target HDR video 54 may be generated that includes the backward HDR frame image g2(0) as the HDR frame image g(0) of the first frame. For the other frames, a target HDR video 54 may be generated that includes the forward HDR frame image g1(t) as the HDR frame image g(t) of the frame.

[0093] As described above, according to this embodiment, it is possible to ensure the temporal consistency of the generated target HDR video 54 while accurately widening the dynamic range of the first frame included in the target HDR video 54.

[0094] Furthermore, in this embodiment, different machine learning models are not used in the forward generation process and the backward generation process, but the machine learning model 30 used in the forward generation process is also used in the backward generation process.

[0095] In video expression, the flow of time can be either forward or backward. Therefore, there is no particular problem in using the machine learning model 30 used in the forward generation process for the backward generation process. In particular, as described above, by having the machine learning model 30 learn videos obtained by playing a given HDR video in reverse, accurate HDR videos can be generated using the common machine learning model 30 in both the forward generation process and the backward generation process.

[0096] As such, in this embodiment, there is no need to prepare separate machine learning models for the forward generation process and the backward generation process, and therefore, according to this embodiment, the effort required for training the machine learning model 30 can be reduced.

[0097] Furthermore, in this embodiment, by referencing the correspondence data, the learning input SDR video 22 that is associated with the teacher HDR video 24 is generated. Therefore, according to this embodiment, it is possible to uniformly and appropriately widen the dynamic range of various SDR images.

[0098] The functions of the image processing device 10 according to this embodiment and the processes executed by the image processing device 10 according to this embodiment will be further described below.

[0099] Fig. 9 is a functional block diagram showing an example of functions implemented in the image processing device 10 according to this embodiment. Note that the image processing device 10 according to this embodiment does not need to implement all of the functions shown in Fig. 9, and functions other than the functions shown in Fig. 9 may also be implemented.

[0100] As shown in FIG. 9, the image processing device 10 according to this embodiment functionally includes, for example, a machine learning model 30, a teacher HDR video storage unit 60, a correspondence data storage unit 62, a training data generation unit 64, a training data storage unit 66, a learning unit 68, a target SDR video acquisition unit 70, a forward generation processing execution unit 72, a backward generation processing execution unit 74, and a target HDR video generation unit 76.

[0101] The machine learning model 30 is implemented mainly using the processor 12 and the storage unit 14. The teacher HDR video storage unit 60, the correspondence data storage unit 62, and the training data storage unit 66 are implemented mainly using the storage unit 14. The training data generation unit 64, the learning unit 68, the target SDR video acquisition unit 70, the forward generation process execution unit 72, the backward generation process execution unit 74, and the target HDR video generation unit 76 are implemented mainly using the processor 12.

[0102] The image processing device 10 according to this embodiment serves both as a learning device that executes learning of a machine learning model 30 that, in response to input SDR images, outputs HDR images by expanding the dynamic range of the input SDR images, and as a video generation device that uses the trained machine learning model 30 (trained model) to generate HDR video by converting SDR video to HDR based on the input SDR video. In the example of FIG. 9 , the machine learning model 30, the teacher HDR video storage unit 60, the correspondence data storage unit 62, the training data generation unit 64, the training data storage unit 66, and the learning unit 68 function as the learning device. The machine learning model 30, the target SDR video acquisition unit 70, the forward generation process execution unit 72, the backward generation process execution unit 74, and the target HDR video generation unit 76 function as the video generation device.

[0103] The above functions may be implemented by executing a program including instructions corresponding to the above functions, which is installed in the image processing device 10, which is a computer, on the processor 12. This program may be supplied to the processor 12 via a computer-readable information storage medium such as an optical disk, a magnetic disk, a magnetic tape, a magneto-optical disk, or a flash memory, or via the Internet, for example.

[0104] In this embodiment, the machine learning model 30 is, for example, a machine learning model that generates an HDR frame image associated with an SDR frame image based on an SDR frame image and an HDR frame image of a frame adjacent to the SDR frame image. Like the machine learning model 30 described in the example above, the machine learning model 30 may be a machine learning model that generates an HDR frame image associated with an SDR frame image based on an SDR frame image, an SDR frame image of a frame adjacent to the SDR frame image, and an HDR frame image of the adjacent frame. In addition, in this embodiment, the machine learning model 30 is, for example, a machine learning model that, in response to an input SDR image, outputs an HDR image with an expanded dynamic range of the image.

[0105] The machine learning model 30 may not be input with the SDR frame image a(t-1) shown in FIG. 4, the SDR frame image x(t-1) shown in FIG. 7, or the SDR frame image x(1) shown in FIG. 8. In this case, the machine learning model 30 may not include the first connection block 32. The feature extraction block 34 may be input with the HDR frame image c(t-1) shown in FIG. 4, the forward HDR frame image g1(t-1) shown in FIG. 7, or the forward HDR frame image g1(1) shown in FIG. 8.

[0106] In this embodiment, the teacher HDR moving image storage unit 60 stores, for example, a plurality of the teacher HDR moving images 24 described above.

[0107] In this embodiment, the correspondence data storage unit 62 stores correspondence data in which, for example, luminance values ​​in HDR and luminance values ​​in SDR are associated with each other. As described above, the correspondence data may be a lookup table such as a one-dimensional lookup table.

[0108] Furthermore, as shown in FIG. 3, the correspondence data may be data in which a luminance value equal to or greater than a predetermined value P1 in HDR is associated with a saturation value Q1 in SDR.

[0109] In this embodiment, the training data generation unit 64 generates, for example, the above-described training data 20. The training data generation unit 64 may, for example, refer to the correspondence data stored in the correspondence data storage unit 62 to generate, based on an HDR image, an SDR image associated with the HDR image. For example, the training data generation unit 64 may, based on a teacher HDR video 24, generate a training input SDR video 22 associated with the teacher HDR video 24 by referencing the correspondence data stored in the correspondence data storage unit 62. Then, the training data generation unit 64 may generate training data 20 including the teacher HDR video 24 and the training input SDR video 22 associated with the teacher HDR video 24.

[0110] The training data generation unit 64 may then store, in the training data storage unit 66, a plurality of sets of training data 20 each associated with a plurality of training HDR videos 24.

[0111] In this embodiment, the training data storage unit 66 stores, for example, a plurality of training data 20 generated by the training data generation unit 64.

[0112] In the present embodiment, the learning unit 68 performs training of the machine learning model 30 using, for example, SDR images and HDR images associated with the SDR images. For example, as described above, the learning unit 68 may perform training of the machine learning model 30 using the training input SDR video 22 included in the training data 20 and the teacher HDR video 24 included in the training data 20. Alternatively, the learning unit 68 may perform training of the machine learning model 30 using the reference HDR video 50 generated based on the training input SDR video 22 included in the training data 20 and the teacher HDR video 24 included in the training data 20.

[0113] In this embodiment, the target SDR video acquisition unit 70 acquires, for example, the above-mentioned target SDR video 52, which is an SDR video to be input to the trained machine learning model 30 (trained model).

[0114] In this embodiment, for example, the forward generation process execution unit 72 uses a trained machine learning model 30 to perform forward generation process to generate an HDR frame image associated with an SDR frame image based on an SDR frame image and an HDR frame image of the frame immediately preceding the SDR frame image generated by the machine learning model 30.

[0115] For example, the forward generation process execution unit 72 uses the trained machine learning model 30 to generate a forward HDR frame image g1(t) based on an SDR frame image x(t) and a forward HDR frame image g1(t-1).

[0116] As described above, the forward generation process execution unit 72 may generate a forward HDR frame image g1(t) based on the SDR frame image x(t), the SDR frame image x(t-1), and the forward HDR frame image g1(t-1) using the trained machine learning model 30. Note that the SDR frame image x(t-1) does not have to be included in the input to the trained machine learning model 30.

[0117] Furthermore, the forward generation process execution unit 72 does not have to generate forward HDR frame images g1 for all SDR frame images x.

[0118] In this embodiment, for example, the backward generation process execution unit 74 uses the trained machine learning model 30 to perform backward generation process for at least the first frame of an SDR frame image, in which an HDR frame image corresponding to the SDR frame image is generated based on the SDR frame image and the HDR frame image of the frame immediately following the SDR frame image generated by the machine learning model 30.

[0119] For example, the backward generation process execution unit 74 uses the trained machine learning model 30 to generate a backward HDR frame image g2(0) based on the SDR frame image x(0) and the forward HDR frame image g1(1).

[0120] For example, as described above, the backward generation process execution unit 74 may generate a backward HDR frame image g2(0) that is associated with the SDR frame image x(0), based on the SDR frame image x(0), the SDR frame image x(1), and the forward HDR frame image g1(1), using the trained machine learning model 30. Note that the SDR frame image x(1) does not have to be included in the input to the trained machine learning model 30.

[0121] In this embodiment, the target HDR video generation unit 76 generates the target HDR video 54 based on, for example, a forward HDR frame image g1 generated in the forward generation process and a backward HDR frame image g2 generated in the backward generation process. As described above, the target HDR video generation unit 76 may generate an HDR video including the backward HDR frame image g2 of the first frame generated in the backward generation process and the forward HDR frame images g1 of the second and subsequent frames generated in the forward generation process.

[0122] Furthermore, as shown in FIG. 10 , in this embodiment, the reverse generation process execution unit 74 may generate reverse HDR frame images g2 for frames other than the first frame. For example, the reverse generation process execution unit 74 may generate reverse HDR frame images g2 (g2(0), g2(1), g2(2), g2(3), ...) for all frames. Note that the numbers in parentheses correspond to the frame numbers; for example, the reverse HDR frame image g2 of the kth frame will be expressed as reverse HDR frame image g2(k-1).

[0123] Furthermore, the reverse generation process execution unit 74 may generate an HDR frame image associated with an SDR frame image based on an SDR frame image and an HDR frame image of the frame immediately following the SDR frame image generated by the reverse generation process.

[0124] Here, for example, the SDR frame images x may be input to the machine learning model 30 sequentially in reverse order of frame number, starting from the last SDR frame image x (assumed to be x(N)).

[0125] 11, in the backward generation process, not only the SDR frame image x(t) included in the target SDR video 52 but also the SDR frame image x(t+1) may be input to the machine learning model 30. Furthermore, the backward HDR frame image g2(t+1) that is the output of the immediately preceding frame may also be input to the machine learning model 30. In this way, the backward HDR frame image g2(t+1) output from the machine learning model 30 by the backward generation process may be used as input to the machine learning model 30 in the backward generation process.

[0126] The first connection block 32 may then connect the SDR frame image x(t+1) and the reverse HDR frame image g2(t+1) to generate first intermediate data 42. The first intermediate data 42 may then be input to the feature extraction block 34.

[0127] The feature extraction block 34 and the resizing block 36 may then perform the same processing as described above in training the machine learning model 30 to generate the enlarged feature map 46.

[0128] The second connection block 38 may then generate second intermediate data 48 by connecting the SDR frame image x(t) and the enlarged feature map 46. The second intermediate data 48 may then be input to the image generation block 40.

[0129] Then, the image generation block 40 may output the reverse HDR frame image g2(t) in response to input of the second intermediate data 48. In this manner, the reverse HDR frame image g2(t) may be generated.

[0130] In this case, when generating the last frame reverse HDR frame image g2 (referred to as g2(N)) (initial reverse generation process), in addition to the SDR frame image x(N), a dummy image SDR frame image x(N+1) and a dummy image reverse HDR frame image g2(N+1) may be input to the machine learning model 30. Here, an image in which all pixels have the same luminance value may be used as the dummy image. For example, an image in which all pixels are white pixels or an image in which all pixels are black pixels may be used as the dummy image.

[0131] It should be noted that the SDR frame image x(t+1) does not have to be included in the input to the trained machine learning model 30.

[0132] Furthermore, in this embodiment, the target HDR moving image generation unit 76 may generate an HDR frame image g(t) included in the target HDR moving image 54 based on a forward HDR frame image g1(t) and a reverse HDR frame image g2(t) for at least one frame.

[0133] For example, the target HDR moving image generation unit 76 may generate a target HDR moving image 54 in which, for at least one frame, the weighted average of the luminance value of a pixel included in the forward HDR frame image g1(t) and the luminance value of the pixel included in the reverse HDR frame image g2(t) is set to the luminance value of the pixel included in the HDR frame image g(t).

[0134] For example, the luminance value of a pixel included in the forward HDR frame image g1(t) and the luminance value of the pixel included in the reverse HDR frame image g2(t) may be weighted by a predetermined weight (for example, 2:1) and the weighted average may be set as the luminance value of the pixel included in the HDR frame image g(t).

[0135] Alternatively, the smaller the frame number, the heavier the weight of the luminance value of the backward HDR frame image g2(t), and the larger the frame number, the heavier the weight of the luminance value of the forward HDR frame image g1(t). In this case, the weight of the luminance value of the forward HDR frame image g1(t) for the first frame may be 0. Also, the weight of the luminance value of the backward HDR frame image g2(t) for the last frame may be 0.

[0136] An example of the flow of the learning process performed by the image processing device 10 according to this embodiment will be described below with reference to the flow diagram shown in Fig. 12. In this processing example, it is assumed that a plurality of given teacher HDR videos 24 are stored in advance in the teacher HDR video storage unit 60.

[0137] First, the training data generation unit 64 acquires one training HDR video 24 on which the processes shown in S102 to S104 have not yet been performed (S101).

[0138] Then, the training data generation unit 64 refers to the correspondence data stored in the correspondence data storage unit 62 to identify the luminance value in SDR of each pixel included in the HDR frame image b of the teacher HDR moving image 24 acquired by the process shown in S101 (S102).

[0139] Then, the training data generation unit 64 generates the learning input SDR video sequence 22 based on the luminance values ​​identified in the process shown in S102 (S103).

[0140] The training data generation unit 64 then generates training data 20 including the learning input SDR video 22 generated in the process shown in S103 and the teacher HDR video 24 acquired in the process shown in S101. The training data generation unit 64 then stores the training data 20 in the training data storage unit 66 (S104).

[0141] Then, the training data generation unit 64 checks whether or not the processes shown in S101 to S104 have been executed for all the teacher HDR videos 24 stored in the teacher HDR video storage unit 60 (S105).

[0142] If it is confirmed that the processes shown in S101 to S104 have not been executed for all the teacher HDR moving images 24 (S105: N), the process returns to S101.

[0143] If it is confirmed that the processes shown in S101 to S104 have been performed for all the teacher HDR videos 24 (S105: Y), the learning unit 68 acquires one training data 20 for which the process shown in S107 has not yet been performed (S106).

[0144] Then, the learning unit 68 uses the training data 20 acquired in the process shown in S106 to perform learning of the machine learning model 30 (S107).

[0145] Then, the learning unit 68 checks whether or not the process shown in S107 has been executed for all of the training data 20 stored in the training data storage unit 66 (S108).

[0146] If it is confirmed that the process shown in S107 has not been executed for all the training data 20 (S107: N), the process returns to the process shown in S106.

[0147] If it is confirmed that the process shown in S107 has been executed for all the training data 20 (S107: Y), the process shown in this processing example is ended.

[0148] Hereinafter, an example of the flow of the process of generating the target HDR video 54 performed by the image processing device 10 according to this embodiment will be described with reference to the flow diagram shown in FIG.

[0149] First, the target SDR video acquisition unit 70 acquires the target SDR video 52 (S201).

[0150] Then, the forward generation process execution unit 72 executes the forward generation process on the target SDR video 52 acquired in the process shown in S201 to generate a forward HDR frame image g1 for at least one frame (S202).

[0151] Then, the reverse generation process execution unit 74 executes the reverse generation process on the target SDR video 52 acquired in the process shown in S201 to generate a reverse HDR frame image g2 for at least the first frame (S203).

[0152] Then, the target HDR video generation unit 76 generates the target HDR video 54 based on the forward HDR frame image g1 generated in the process shown in S202 and the backward HDR frame image g2 generated in the process shown in S203 (S204). Then, the process shown in this processing example ends.

[0153] The present invention is not limited to the above-described embodiment.

[0154] For example, the scope of application of the present invention is not limited to SDR or HDR.

[0155] For example, the present invention is generally applicable to an image processing device 10 that performs learning of a machine learning model 30 that, in response to an input of a first type of image, which is an image of a first dynamic range (not limited to SDR), outputs a second type of image, which is an image of a second dynamic range (not limited to HDR) in which the dynamic range of the first type of image is expanded.

[0156] In this case, the correspondence data storage unit 62 may store correspondence data in which luminance values ​​in the second dynamic range correspond to luminance values ​​in the first dynamic range. Here, the correspondence data may be data in which luminance values ​​equal to or greater than a predetermined value in the second dynamic range correspond to saturation values ​​in the first dynamic range. The correspondence data may also be a lookup table, such as a one-dimensional lookup table.

[0157] Then, the training data generation unit 64 may refer to the correspondence data to generate, based on the second type image, a first type image that is associated with the second type image.

[0158] Then, the learning unit 68 may perform learning of the machine learning model 30 using the first-type image and the second-type image associated with the first-type image.

[0159] Furthermore, the present invention is generally applicable to an image processing device 10 that uses a trained machine learning model 30 (trained model) to generate, based on a first type of moving image that is a moving image of a first dynamic range (not limited to SDR), a second type of moving image that is a moving image of a second dynamic range (not limited to HDR) that expands the dynamic range of the first type of moving image.

[0160] In this case, the trained machine learning model 30 may be a machine learning model that generates a second type of frame image that corresponds to a first type of frame image, based on the first type of frame image that is a frame image included in a first type of moving image and a second type of frame image that is a frame image of a second dynamic range of a frame adjacent to the first type of frame image.

[0161] The forward generation process execution unit 72 may then use the trained machine learning model 30 to execute a forward generation process to generate a second type of frame image that corresponds to the first type of frame image based on the first type of frame image and a second type of frame image that is the frame immediately before the first type of frame image generated by the machine learning model 30.

[0162] Then, the backward generation process execution unit 74 may use the trained machine learning model 30 to perform a backward generation process for at least the first frame of a first type frame image, based on the first type frame image and a second type frame image of the frame immediately following the first type frame image generated by the machine learning model 30, to generate a second type frame image that corresponds to the first type frame image.

[0163] Here, the backward generation process execution unit 74 may generate a second type of frame image corresponding to the first type of frame image based on the first type of frame image and a second type of frame image that is the frame immediately following the first type of frame image generated in the forward generation process.

[0164] Alternatively, the backward generation process execution unit 74 may generate a second type of frame image corresponding to a first type of frame image based on the first type of frame image and a second type of frame image that is the frame immediately following the first type of frame image generated by the backward generation process.

[0165] Then, the target HDR video generation unit 76 may generate a second type of video based on the second type of frame images generated in the forward generation process and the second type of frame images generated in the backward generation process.

[0166] Furthermore, the trained machine learning model 30 may be a machine learning model that generates a second type of frame image that corresponds to a first type of frame image based on the first type of frame image, a first type of frame image of a frame adjacent to the first type of frame image, and a second type of frame image of the adjacent frame.

[0167] In this case, the forward generation process execution unit 72 may generate a second type of frame image corresponding to the first type of frame image based on the first type of frame image, the first type of frame image of the frame immediately preceding the first type of frame image, and the second type of frame image of the immediately preceding frame generated by the machine learning model 30.

[0168] Then, the backward generation process execution unit 74 may generate a second type of frame image that corresponds to the first type of frame image based on the first type of frame image, the first type of frame image of the frame immediately following the first type of frame image, and the second type of frame image of the frame immediately following the first type of frame image that is generated by the machine learning model 30.

[0169] In addition, the target HDR moving image generation unit 76 may generate a second type of moving image including a second type of frame image of the first frame generated in the backward generation process and a second type of frame image of the second or subsequent frames generated in the forward generation process.

[0170] Alternatively, the target HDR video generation unit 76 may generate, for at least one frame, a frame image of the frame included in the second type of video based on the second type frame image of the frame generated by the forward generation process and the second type frame image of the frame generated by the backward generation process.

[0171] In this case, the target HDR moving image generation unit 76 may generate, for at least one frame, a second type of moving image in which the weighted average of the luminance value of a pixel included in the second type of frame image of the frame generated by the forward generation process and the luminance value of the pixel included in the second type of frame image of the frame generated by the backward generation process is set to the luminance value of the pixel included in the frame image of the frame.

[0172] Furthermore, for example, by referring to correspondence data, an SDR still image corresponding to an HDR still image may be generated based on the HDR still image. Furthermore, the machine learning model 30 does not need to incorporate an RNN function. Then, learning of the machine learning model 30 may be performed using the HDR still image and the SDR still image.

[0173] Furthermore, for example, the image processing device 10 does not have to have a function as a learning device that executes learning of the machine learning model 30 that outputs an HDR image in response to an input of an SDR image. Furthermore, the image processing device 10 does not have to have a function as a moving image generation device that uses a trained machine learning model 30 (trained model) to generate an HDR moving image based on an SDR moving image by converting the SDR moving image into an HDR moving image.

[0174] Furthermore, the specific character strings and numerical values ​​described above and the specific character strings and numerical values ​​in the drawings are examples, and the present invention is not limited to these character strings and numerical values.

Claims

1. A moving image generating device that generates, based on a first type moving image that is a moving image with a first dynamic range, a second type moving image that is a moving image with a second dynamic range obtained by widening the dynamic range of the first type moving image, using a trained model, the trained model is a trained model that generates a second type frame image associated with a first type frame image, based on a first type frame image that is a frame image included in the first type moving image and a second type frame image that is a frame image of the second dynamic range that is adjacent to the first type frame image, The moving image generating device a forward generation processing execution unit that executes a forward generation processing to generate the second type frame image associated with the first type frame image based on the first type frame image and the second type frame image that is generated by the trained model and is a frame immediately before the first type frame image; and a backward generation processing execution unit that executes a backward generation processing for generating, for at least the first type frame image of an initial frame, the second type frame image associated with the first type frame image based on the first type frame image and the second type frame image of the frame immediately following the first type frame image generated by the trained model, using the trained model; a moving image generating unit that generates the second type moving image based on the second type frame images generated in the forward direction generation process and the second type frame images generated in the backward direction generation process; A moving image generating device comprising:

2. the trained model is a trained model that generates the second type frame image associated with the first type frame image based on the first type frame image, the first type frame image of a frame adjacent to the first type frame image, and the second type frame image of the adjacent frame, the forward generation process execution unit generates the second type frame image associated with the first type frame image based on the first type frame image, the first type frame image of a frame immediately before the first type frame image, and the second type frame image of the frame immediately before the first type frame image generated by the trained model; the backward generation process executing unit generates the second type frame image associated with the first type frame image based on the first type frame image, the first type frame image of a frame immediately following the first type frame image, and the second type frame image of the frame immediately following the first type frame image generated by the trained model; 2. The moving image generating device according to claim 1.

3. the backward generation processing execution unit generates the second-type frame image associated with the first-type frame image based on the first-type frame image and the second-type frame image that is a frame immediately following the first-type frame image generated in the forward generation processing; 3. The moving image generating device according to claim 1 or 2.

4. the backward generation processing execution unit generates the second-type frame image associated with the first-type frame image based on the first-type frame image and the second-type frame image that is a frame immediately following the first-type frame image generated by the backward generation processing; 3. The moving image generating device according to claim 1 or 2.

5. the moving image generation unit generates the second type moving image including the second type frame image of a first frame generated in the backward generation process and the second type frame images of second and subsequent frames generated in the forward generation process.

5. The moving image generating device according to claim 1, wherein the moving image generating device is a video image generating device.

6. the moving image generation unit generates, for at least one frame, a frame image of the frame included in the second type moving image based on the second type frame image of the frame generated in the forward generation process and the second type frame image of the frame generated in the backward generation process; 5. The moving image generating device according to claim 1, wherein the moving image generating device is a video image generating device.

7. the moving image generation unit generates, for at least one frame, the second-type moving image in which a weighted average of a luminance value of a pixel included in the second-type frame image of the frame generated by the forward generation process and a luminance value of the pixel included in the second-type frame image of the frame generated by the backward generation process is set to the luminance value of the pixel included in the frame image of the frame.

7. The moving image generating device according to claim 6.

8. the first dynamic range is a standard dynamic range (SDR); the second dynamic range is a high dynamic range (HDR); 8. The moving image generating device according to claim 1,

9. A moving image generating method for generating, based on a first type moving image that is a moving image with a first dynamic range, a second type moving image that is a moving image with a second dynamic range obtained by widening the dynamic range of the first type moving image, using a trained model, the trained model is a trained model that generates a second type frame image associated with a first type frame image, based on a first type frame image that is a frame image included in the first type moving image and a second type frame image that is a frame image of the second dynamic range that is adjacent to the first type frame image, The moving image generating method includes: a step of executing a forward generation process using the trained model to generate the second type frame image associated with the first type frame image based on the first type frame image and the second type frame image that is a frame immediately before the first type frame image generated by the trained model; using the trained model, for at least the first type frame image of an initial frame, based on the first type frame image and the second type frame image of the frame immediately following the first type frame image generated by the trained model, performing a backward generation process to generate the second type frame image associated with the first type frame image; generating the second type of moving image based on the second type of frame images generated in the forward generation process and the second type of frame images generated in the backward generation process; A moving image generating method comprising:

10. A program that causes a computer to generate, based on a first type of moving image that is a moving image with a first dynamic range, a second type of moving image that is a moving image with a second dynamic range obtained by widening the dynamic range of the first type of moving image, using a trained model, the trained model is a trained model that generates a second type frame image associated with a first type frame image, based on a first type frame image that is a frame image included in the first type moving image and a second type frame image that is a frame image of the second dynamic range that is adjacent to the first type frame image, The program causes the computer to: a step of executing a forward generation process using the trained model to generate the second-type frame image associated with the first-type frame image based on the first-type frame image and the second-type frame image that is generated by the trained model and is a frame immediately before the first-type frame image; a step of executing a backward generation process using the trained model to generate, for at least a first-type frame image of the first type, a second-type frame image associated with the first-type frame image based on the first-type frame image and the second-type frame image of the frame immediately following the first-type frame image generated by the trained model; generating the second type of moving image based on the second type of frame images generated in the forward direction generation process and the second type of frame images generated in the backward direction generation process; A program characterized by executing the following.

Citation Information

Patent Citations

  • Generating a high dynamic range image from a low dynamic range image

    JP2013539610A

  • Real-time reconfiguration of single-layer backward compatible codecs

    JP2019530309A

  • Tone mapping processing, HDR video conversion method by automatic adjustment and update of tone mapping parameter, and device of the same

    JP2020017079A

  • Information processing device and information processing method

    JP2020154609A

  • Efficient End-to-End Single-Layer Inverse Display Management Coding

    JP2020524446A