Information processing device, information processing method, and computer-readable non-transitory storage medium

The information processing apparatus addresses the issue of decreased DNN inference accuracy due to rendering condition changes by using a specialization coefficient database, feature amount conversion, motion compensation, and a DNN to maintain accurate video processing.

WO2025121063A1PCT designated stage expired Publication Date: 2025-06-12SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/039439
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-11-06
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

When rendering conditions change between frames, the inference accuracy of a Deep Neural Network (DNN) may decrease due to the input of intermediate feature amounts that the switched specialization coefficient has not learned, leading to issues like color shift and resolution degradation in the output image.

Method used

An information processing apparatus that includes a specialization coefficient database to switch coefficients based on rendering condition changes, a feature amount conversion unit to convert intermediate feature amounts from before the switch to match the new coefficient, a motion compensation unit to adjust the converted feature amounts, and a DNN that performs inference using the corrected history frame.

Benefits of technology

This configuration helps maintain high inference accuracy of the DNN even when rendering conditions switch, reducing the likelihood of color shift and resolution degradation in the output video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024039439_12062025_PF_FP_ABST
    Figure JP2024039439_12062025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device includes a specialization coefficient database, a feature amount conversion unit, a motion compensation unit, and a DNN. The specialization coefficient database switches a specialization coefficient in accordance with switching of the rendering condition from the past frame to the input current frame. The feature amount conversion unit converts the intermediate feature amount related to the past frame having been inferred using the specialization coefficient before the rendering condition is switched into a correction feature amount estimated to be inferred by the specialization coefficient after the rendering condition is switched. The motion compensation unit performs motion compensation on the correction feature amount to acquire a correction history frame. The DNN performs inference regarding the input current frame on the basis of the correction history frame and the previous input current frame.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and computer-readable non-transitory storage medium

[0001] The present invention relates to an information processing device, an information processing method, and a computer-readable non-transitory storage medium.

[0002] Deep neural network (DNN) processing for video generally uses a recurrent neural network (RNN) structure with high time stability. This is because the RNN structure can correlate the current frame (present information) with the history frames (past information), making it easier to maintain the consistency of the output frame image.

[0003] It is known that highly accurate estimation results can be obtained by using features (intermediate features) output from the intermediate layer of a DNN as history frames. The features extracted by the DNN can only be determined after the DNN has been trained. Even under the same training conditions, the resulting features will differ depending on the randomness of the training process.

[0004] Japanese Patent Application Laid-Open No. 2022-008037

[0005] In general, the inference accuracy of a DNN is higher when the characteristics of the input data at the time of inference and the input data at the time of learning are closer. If the rendering conditions (resolution, characteristics of the video scene to be specialized, etc.) of the input image used at the time of inference are known, it is desirable to use coefficients (specialized coefficients) learned using data specialized for those rendering conditions.

[0006] When rendering conditions change between frames, it is desirable to also change the DNN specialization coefficients in accordance with the change in rendering conditions. For example, in game consoles, the processing load of the renderer fluctuates, and video delays may occur when the load is high. Therefore, the resolution of the rendered image (rendering resolution) is changed according to the processing load to maintain the frame rate. This technology is called Dynamic Resolution (DR).

[0007] However, when the specialization coefficient is switched, the intermediate features output before the specialization coefficient was switched are input to the DNN after the specialization coefficient was switched. The input intermediate features are features that the specialization coefficient after the switch has never handled during learning, i.e., features that have never been learned. This can result in a decrease in the inference accuracy of the DNN, and problems such as color shift and reduced resolution can occur in the output image.

[0008] Therefore, the present disclosure proposes an information processing device, an information processing method, and a computer-readable non-transitory storage medium that are less likely to reduce the inference accuracy of a DNN when rendering conditions are switched.

[0009] According to the present disclosure, there is provided an information processing device including: a specialized coefficient database that switches specialized coefficients in accordance with switching of rendering conditions from a past frame to an input current frame; a feature conversion unit that converts intermediate features related to the past frame estimated using the specialized coefficients before the switching of the rendering conditions into corrected features estimated to be estimated using the specialized coefficients after the switching of the rendering conditions; a motion compensation unit that performs motion compensation on the corrected features to obtain corrected history frames; and a DNN that performs inference related to the input current frame based on the corrected history frames and the input current frame. Also according to the present disclosure, there is provided an information processing method in which information processing of the information processing device is executed by a computer, and a program that causes a computer to realize the information processing of the information processing device.

[0010] FIG. 1 is an explanatory diagram of DR. FIG. 2 is an explanatory diagram of specialization coefficients of DNN. FIG. 3 is an explanatory diagram of the operation of DNN in video processing. FIG. 4 is an explanatory diagram of the use of specialization coefficients in a scene where DR is applied. FIG. 5 is a diagram illustrating an example of the configuration of an information processing device. FIG. 6 is a diagram illustrating an example of conversion of DNN output in accordance with switching of specialization coefficients. FIG. 7 is a diagram illustrating an example of a method for calculating conversion parameters. FIG. 8 is a diagram illustrating intermediate feature amounts. FIG. 9 is a diagram illustrating optimization processing of conversion parameters. FIG. 10 is a diagram illustrating a generalized expression of a conversion model. FIG. 11 is a diagram illustrating an example of a processing flow for performing inference processing. FIG. 12 is a diagram illustrating an example of a processing flow for performing inference processing. FIG. 13 is a diagram illustrating an example of a processing flow for calculating conversion parameters. FIG. 14 is a diagram illustrating an example of the hardware configuration of an information processing device.

[0011] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0012] The explanation will be given in the following order: [1. Dynamic Resolution (DR)] [2. Specialization coefficients of DNN] [3. DNN in video processing] [4. Use of specialization coefficients in DR application scenes] [5. Example system configuration of information processing device] [6. Conversion of DNN output (intermediate feature amount) in accordance with switching of specialization coefficients] [7. Calculation of conversion parameters] [8. Processing flow] [9. Example hardware configuration] [10. Effects]

[0013] In the present disclosure, the specialization coefficient D of the DNN is adjusted according to the switching of the rendering conditions between frames. R In the present disclosure, the specialization coefficient D before the rendering condition is switched (see FIG. 5) is used. R The intermediate feature inferred in (1) is used as the specialization coefficient D after the rendering conditions are switched. R The corrected features obtained by the conversion are then motion-compensated and used as recurrent data, which prevents degradation of inference accuracy when the rendering conditions are changed.

[0014] Specialization coefficient D Ris learned as a coefficient specialized for specific rendering conditions. Rendering conditions that can be specialized include rendering resolution and scene characteristics for which the rendered image is specialized (such as an explosion scene). In the following explanation, the scene in which the renderer is used is real-time rendering of a game engine, and the specialized coefficient D R An example will be described in which σ is a coefficient learned for each rendering resolution.

[0015] [1. Dynamic Resolution (DR)] FIG. 1 is an explanatory diagram of DR.

[0016] In game consoles, the processing load of the renderer fluctuates, and video delays can occur when the load is high. Examples of high-load rendering include explosion effects, scenes with a high number of polygons, communication synchronization waits, and ray tracing execution. In the example of Figure 1, a 4K rendered image is displayed at 60 FPS when the load is low, but when the load is high, the rendered image under the same conditions can only be displayed at 15 FPS (video delay).

[0017] To avoid such video delays, the rendering resolution is changed according to the load, and the frame rate is maintained. For example, when the load is high, the output resolution of the renderer is lowered, and then an upscaling process is applied to match the display resolution. This technique is called Dynamic Resolution (DR). In the example of FIG. 1, when the load is high, an image rendered at 2K is upscaled by two times using a scaler, resulting in a 4K image being displayed. Because the rendering load (the calculation cost of the drawing process) is reduced, the frame rate can be maintained at 60 FPS.

[0018] [2. DNN specialization coefficient] Figure 2 shows the DNN specialization coefficient D R FIG.

[0019] The inference accuracy of DNN is higher when the characteristics of the input data at the time of inference and the input data at the time of learning are closer. If the rendering conditions of the input image used at the time of inference are known, the coefficients learned using data specialized for those rendering conditions (specialized coefficients D RIn a DR application scenario, the rendering resolution is adaptively switched depending on the processing load. In a situation where the rendering resolution changes, such as 1K, 2K, or 4K, the specialization coefficient D R By applying this, high-quality inference results can be obtained.

[0020] 3. DNN in Video Processing FIG. 3 is a diagram illustrating the operation of DNN in video processing.

[0021] DNN processing for video generally uses an RNN structure with high temporal stability. The RNN structure can correlate the current frame with history frames, making it easier to maintain image consistency in the output frames. While DNNs have very high inference performance (for example, sharpening effects in the case of super-resolution), they are affected by even small changes in the input data, which can change the output results. Therefore, in video processing, flickering is more likely to occur than with simple filter processing. To increase temporal stability and reduce flickering, an RNN structure is generally implemented that inputs the current frame and history frames simultaneously, enabling learning that includes temporal correlation (similarity between the current frame and history frames).

[0022] Note that the current frame refers to the image of the frame (current frame) that is the subject of estimation by the DNN. The history frame refers to an image (inference image) obtained by inputting an image of a past frame into the DNN, or intermediate features extracted from an image of a past frame. The past frame refers to a frame that is one or more frames before the current frame. In the present disclosure, for example, the frame that is one frame before the current frame (the most recent frame) is used as the past frame. Hereinafter, the current frame may be referred to as the "input current frame," and the history frame that is the inference result of the past frame may be referred to as the "inference history frame." Furthermore, the inference image inferred by the DNN and output as an RGB image may be referred to as the "output image."

[0023] Intermediate features refer to information on features output from the intermediate layer of a DNN when an image of a past frame is input to the DNN. It is known that using intermediate features obtained during inference as an inference history frame can produce more accurate estimation results than using an inference image (RGB image) of a past frame as an inference history frame. In the present disclosure, intermediate features obtained during inference are used as the inference history frame.

[0024] [4. Use of specialization coefficient in DR application scene] Figure 4 shows the specialization coefficient D R FIG.

[0025] As mentioned above, in the application scene of DR, the specialization coefficient D R In the example of FIG. 4, 4K rendering is performed at time t=N. The coefficients of the DNN are the specialized coefficients D R (4K specialization coefficient) is used.

[0026] At time t=N+1, the rendering resolution is switched depending on the processing load. Rendering is performed at 2K resolution, and the DNN coefficients are the specialized coefficients D R (2K specialization coefficient). The 2K rendered image is upscaled by 2 times and then input to the DNN.

[0027] However, the inference history frame input to the DNN with 2K specializing coefficients at time t=N+1 is an intermediate feature inferred using 4K specializing coefficients at time t=N. This intermediate feature is likely to be a feature with a different nature from that used to learn the 2K specializing coefficients. If an intermediate feature with a completely unanticipated, unlearned nature is input to a DNN whose coefficients have been switched to 2K specializing coefficients, high inference accuracy cannot be obtained, and failures such as color shifts and reduced resolution may occur in the displayed image.

[0028] In the present disclosure, in order to solve such a problem, the specialization coefficient D RThe intermediate feature value output before the switching is used as the specialization coefficient D R The generation of the conversion model CV (see FIG. 9 ) for performing the conversion processing and the execution of the inference processing using the converted intermediate features (corrected features) are realized by the information processing device 1 shown in FIG.

[0029] 5. System Configuration Example of Information Processing Apparatus] FIG. 5 is a diagram showing an example of the configuration of the information processing apparatus 1. As shown in FIG.

[0030] The information processing device 1 includes an input unit 10, a renderer 20, a scaler 30, a DNN 40, a feature conversion unit 50, a motion compensation unit 60, an output unit 70, a conversion parameter calculation unit 80, a feature storage unit 90, and a specialized coefficient database DB C and conversion parameter database DB P It has.

[0031] The input unit 10 receives an operation signal S S is input to the renderer 20. For example, the operation signal S S contains rendering settings and instructions for performing rendering.

[0032] The renderer 20 receives the operation signal S S The renderer 20 performs rendering processing based on the input current frame I C The rendered image is output as the input current frame I C refers to a frame of image in a series of frames such as a video. The block diagram in Figure 5 explains the operation for one frame of a video.

[0033] The renderer 20 obtains the motion vector MV of the subject based on the rendering settings and sends it to the scaler 30. The motion vector MV is a two-dimensional vector that defines the movement of corresponding pixels between frames by the amount of pixel movement (number of pixels). The motion vector MV can be obtained by a general estimation process using a group of consecutive frames. When using a CG renderer, the motion vector MV can be generated and obtained during rendering.

[0034] The renderer 20 obtains the conditions for specializing the rendering (rendering conditions) based on the rendering settings. The renderer 20 stores information about the rendering conditions as rendering information S R The renderer 20 outputs the rendering information S R Specialized coefficient database DB C and conversion parameter database DB P For example, if the rendering is specialized for a specific resolution, the resolution value is acquired as the rendering condition. If the rendering is specialized for a live-action scene or an anime-style scene, scene information is acquired as the rendering condition.

[0035] The scaler 30 scales the input current frame I as needed. C and the number of pixels of the motion vector MV is enlarged. Hereinafter, when it is necessary to distinguish data after the number of pixels has been enlarged by the scaler 30 from data before enlargement, a "'" is added after the symbol of the data.

[0036] The architecture of the DNN 40 accepts a fixed number of pixels, while the input current frame I C In cases where the number of pixels varies depending on the usage, the scaler 30 may perform pixel number expansion processing at a desired magnification ratio so as to obtain a pixel number that matches the specifications of the DNN 40. The pixel number expansion method may be selected arbitrarily, for example, bilinear or bicubic. The motion vector MV is preferably expanded using the nearest neighbor method to avoid the generation of intermediate values ​​due to pixel interpolation, but is not limited to this.

[0037] As mentioned above, the renderer 20 may change the rendering resolution depending on the processing load. In this case, the scaler 30 adjusts the rendering resolution of the input current frame I to match the architecture of the DNN 40. C Scale the

[0038] Specialization coefficient database DB C is the specialization coefficient D of the pre-trained DNN40 R Specialization coefficient database DB Cis the rendering information S R Detects switching of rendering conditions based on the specialization coefficient database DB C is the input current frame I C When a change in the rendering conditions to R and sends it to the DNN 40. C is the specialization coefficient D of DNN40 according to the change of the rendering condition. R Switch between.

[0039] The DNN 40 receives the input current frame I C ' and the correction history frame I Hc MC The DNN 40 performs inference of the current frame based on the inference image of the current frame obtained by inference as an inference current frame. The DNN 40 acquires the feature amounts output from the intermediate layer of the DNN 40 as intermediate feature amounts. The DNN 40 outputs the intermediate feature amounts as an inference history frame I H and outputs it to the feature transform unit 50.

[0040] The DNN 40 calculates the inferred current frame as the output image I O The output unit 70 outputs the output image I O General post-processing such as color conversion and codec is applied to the image. Tasks for the DNN 40 include super-resolution, noise reduction, and style transfer, but the present disclosure is not limited thereto. The DNN 40 can be applied to general video processing using an RNN structure.

[0041] The feature quantity conversion unit 50 converts the most recently acquired inference history frame I H The conversion process is performed using the conversion parameter database DB P The transformation parameters P T The feature conversion unit 50 converts the specialization coefficient D RThe intermediate feature quantity for the past frame inferred using R The intermediate feature (corrected feature I Hc ) to

[0042] As mentioned above, when the rendering conditions are changed, the specialization coefficient D R The specialization coefficient D after switching R Intermediate features applied to the inference history frame I H ) is the specialization coefficient D before switching R The different specialization coefficients D R There is a possibility that intermediate features with different properties are output from a specialization coefficient D R When using the function, even if an intermediate feature value related to the brightness of the input image is output, other specialization coefficients D R When using this method, intermediate features related to the edges of the input image may be output.

[0043] Therefore, the feature conversion unit 50 converts the intermediate feature acquired immediately before the switching of the rendering conditions into the specialization coefficient D R The feature conversion unit 50 converts the converted intermediate feature into a corrected feature I Hc As a result, the intermediate feature amount (inference history frame I) input to the DNN 40 before and after the switching of the rendering conditions is acquired. H ) homogeneity is maintained.

[0044] The feature conversion unit 50 converts the inference history frame I only when the rendering conditions are changed between frames. H When the rendering conditions are maintained between frames, the feature conversion unit 50 converts the inference history frame I H is used as the corrected feature I Hc Output as

[0045] The motion compensation unit 60 receives the corrected feature I from the feature conversion unit 50. Hc The motion compensation unit 60 obtains the corrected feature quantity I based on the motion vector MV'.Hc , and perform motion compensation to obtain the corrected history frame I Hc MC The motion compensation unit 60 performs a common process known as motion compensation to obtain the corrected feature I Hc can be applied to.

[0046] The motion vector MV' has a vector related to the spatial movement of the object between frames. Motion compensation is performed as a process of predicting post-motion data from pre-motion data. The motion compensation unit 60 applies motion compensation to the intermediate feature of the past frame output from the feature conversion unit 50. By motion compensation, the input current frame I C The feature quantity (corrected history frame I Hc MC The DNN 40 calculates the corrected history frame I Hc MC and the input current frame I C ' based on the input current frame I C ' and make inferences about it.

[0047] The feature amount storage unit 90 and the transformation parameter calculation unit 80 calculate the transformation parameter P T The conversion parameter P T In calculating the intermediate feature I, the renderer 20 generates a large number of learning rendering images. For example, the renderer 20 renders the same learning video scene for each rendering condition. The DNN 40 calculates the intermediate feature I for each learning rendering image acquired from the renderer 20. H DR The feature storage unit 90 acquires the intermediate feature I corresponding to each rendering condition. H DR The intermediate feature I H DR is stored in association with the rendering conditions.

[0048] The transformation parameter calculation unit 80 calculates the intermediate feature I H DR A conversion model CV (see FIG. 9) that performs conversion betweenT For example, the transformation process is performed by a matrix operation (transformation matrix), and the transformation parameter P T is obtained as a matrix element. P is the transformation parameter P T and the transformation parameter P T The two rendering conditions to be converted by are stored in association with each other.

[0049] For example, when super-resolution is performed from 2K to 4K, the conversion parameter calculation unit 80 acquires a large number of rendering images generated for each resolution for various scenes. For example, the intermediate feature I obtained by applying the 2K rendering image to the 2K specialization coefficients is H DR is defined as the "2K feature amount," and the intermediate feature amount I obtained by applying the 4K rendering image to the 4K specialization coefficient is defined as H DR The conversion parameter calculation unit 80 applies the 2K feature amount and the 4K feature amount of each scene to the conversion model CV, and calculates the conversion parameter P by a method such as multiple regression analysis. T Optimize the following.

[0050] Conversion parameter database DB P is the rendering information S R Detects a change in rendering conditions based on the conversion parameter database DB P When it is determined that the rendering conditions have been switched, the conversion parameter P associated with the rendering conditions before and after the switching is calculated. T The feature transformation unit 50 sends the transformation parameter database DB P The transformation parameters P obtained from T Applying the transformation model CV to the inference history frame I H The conversion process is performed.

[0051] [6. Conversion of DNN output (intermediate feature amount) in accordance with switching of specialization coefficient] FIG. R 10A and 10B are diagrams illustrating an example of conversion of DNN output in accordance with switching of the

[0052] In the example of FIG. 6, the rendering resolution is switched depending on the processing load of the renderer 20. C acquires information about the rendering resolution as a rendering condition. The renderer 20 performs rendering at a first resolution (for example, 4K) from time t0 to time t. The DNN 40 applies a first specialization coefficient (for example, a 4K specialization coefficient) learned using the rendered image at the first resolution. 4K " means the DNN 40 to which the first specialization coefficient is applied (hereinafter referred to as the first DNN).

[0053] The first DNN receives the input current frame I C as the first resolution image C 4K (t) is obtained. 4K (t) denotes the image obtained by rendering the video scene at a first resolution. The first DNN denotes the first resolution image C of the previous scene. 4K Intermediate feature RNN inferred by itself for (t-1) 4K (t-1) is inferred history frame I H The past scene means a video scene of a past frame. At time t, the specialization coefficient D R The rendering conditions (first resolution) of the learning image of H The rendering conditions (first resolution) of the image that is the source of inference of the first DNN match. Therefore, the inference accuracy of the first DNN is high, and a high-quality output image I is obtained as an inference result. O (t) is obtained.

[0054] The renderer 20 switches the rendering resolution to a second resolution (e.g., 2K) at time (t+1). C is the specialization coefficient D corresponding to the switching of the rendering resolution from the first resolution to the second resolution. R The second specialization coefficient is switched from the first specialization coefficient to the second specialization coefficient D R (e.g., 2K specialization coefficients). 2K" means the DNN 40 to which the second specialization coefficient is applied (hereinafter referred to as the second DNN).

[0055] The second DNN receives the input current frame I C as the second resolution image C 2K (t+1) is acquired. 2K (t+1) denotes the image obtained by rendering the video scene at a second resolution. The second DNN is the first resolution image C of the previous scene. 4K Intermediate feature RNN inferred by the first DNN for (t) 4K (t) is the inference history frame I H At time (t+1), the specialization coefficient D R The rendering conditions (second resolution) of the learning image and the inference history frame I H The rendering conditions (first resolution) of the image that is the source of inference of is different from that of the intermediate feature RNN. 4K If (t) is used as is as recurrent data, the inference accuracy of the second DNN will be low.

[0056] Therefore, the feature transform unit 50 transforms the intermediate feature RNN 4K (t) is the specialization coefficient D after switching R Intermediate feature quantity with properties suitable for (second specialization coefficient) c RNN 2K (t). For example, the intermediate feature RNN 4K (t) is the first resolution image C of the past scene at time t 4K The feature quantity obtained by applying (t) to the first DNN is the second resolution image C. The image obtained by rendering the past scene at the same time t at the second resolution is the second resolution image C. 2K (t), the intermediate feature c RNN 2K (t) is the second resolution image C 2K (t) is defined as a feature that is estimated to be obtained when the second DNN is applied. 4K→2K " represents the conversion model CV.

[0057] The feature transform unit 50 transforms the intermediate feature c RNN 2K (t) is the corrected feature I HcThe feature transformation unit 50 obtains the corrected feature I Hc is output as data for recurrent use. Hc has the same properties as the inference result of the second DNN. At time (t+1), the recurrent data input to the second DNN is a feature quantity of the same quality as the inference result of the second DNN. Therefore, the inference accuracy of the second DNN is increased, and a high-quality output image I is obtained as the inference result. O (t+1) is obtained.

[0058] The rendering resolution remains at the second resolution even after time (t+2). R The second specialization coefficient is maintained at the same value from time (t+2) onwards. R The rendering conditions (second resolution) of the learning image and the inference history frame I H The rendering conditions (second resolution) of the image that is the source of inference of the second DNN match. Therefore, the inference accuracy of the second DNN is high, and a high-quality output image I is obtained as an inference result. O is obtained.

[0059] 7. Calculation of Transformation Parameters FIG. T FIG. 10 is a diagram illustrating an example of a calculation method.

[0060] In the example of Fig. 7, a parameter for conversion between the inferred value of the 4K specialization coefficient and the inferred value of the 2K specialization coefficient is calculated. T In calculating the intermediate feature I, the renderer 20 generates a large number of rendering images for learning. For example, the renderer 20 generates 4K rendering images (4K images) and 2K rendering images (2K images) for each of a large number of video scenes. For the 4K images, the DNN 40 calculates the intermediate feature I using a 4K specialization coefficient. H DR For 2K images, the intermediate feature I H DR Get.

[0061] The scaler 30 may be used as appropriate so that the number of pixels input to the DNN 40 is constant relative to the rendering resolution. A constant number of pixels makes it easier to implement in hardware. In the example of FIG. 7, a 2K image is upscaled by a factor of 2. The intermediate feature I H DR has multiple dimensions (channels) according to the architecture of the DNN 40. In the example of FIG. 7, the intermediate feature I H DR The feature storage unit 90 stores the intermediate feature I H DR is stored in association with the rendering resolution.

[0062] FIG. 8 shows the intermediate feature I H DR FIG. 9 is a diagram illustrating the transformation parameter P T FIG. 10 is a diagram illustrating the optimization process.

[0063] Intermediate feature I H DR indicates the internal feature of the DNN 40. In FIG. C is input to the DNN 40, and the feature z n,l Here, "n" represents the number of dimensions, and "l" represents the layer number. DNN40 is expressed with four layers.

[0064] The final output of DNN40 is z 1,3 ~z 64,3 8, the number of dimensions is 64, but the number of dimensions is determined by the configuration of the DNN 40 and does not necessarily have to be 64.

[0065] In the RNN structure, the inference accuracy is improved by using such multidimensional features rather than recursively using RGB images as input. H The intermediate features recursively used as the multidimensional features (z 1,l ~zn,l In practice, it is desirable to use the feature just before the final layer (the output of the third layer in the example of FIG. 8 ), which is closer to the output, i.e., which is a feature that better represents the training data.

[0066] The conversion parameter calculation unit 80 calculates the intermediate feature RNN inferred using the 4K specialization coefficients. 4K Applying the transformation process to the intermediate feature RNN inferred with 2K specialization coefficients 2K For example, the transformation parameter calculation unit 80 approximates the intermediate feature RNN 4K The feature obtained by applying the conversion model CV to the intermediate feature RNN 2K The conversion parameter calculation unit 80 optimizes the conversion model CV so that it approximates the individual coefficients included in the optimized conversion model CV as conversion parameters P T Obtain as.

[0067] The optimization of the conversion model CV can be performed using a general regression model. FIG. 9 shows an example of multiple regression analysis. In FIG. 9, "c" and "b" indicate coefficients of the conversion model CV. In the example of FIG. 9, the feature has a total of eight dimensions, from 0th to 7th. Note that the number of dimensions of the feature is not limited to this. The number of dimensions of the feature may be nine or more. FIG. 10 is a diagram showing a generalized expression of the conversion model CV. The conversion process when the rendering condition is switched from α to β can be generalized and expressed by the formula in FIG. 10, where m is the total number of pixels and ch is the number of dimensions of the intermediate feature.

[0068] 8. Processing Flow FIGS. 11 and 12 are diagrams showing an example of a processing flow for performing inference processing.

[0069] The scaler 30 receives the input current frame I from the renderer 20. C , the motion vector MV and the rendering information S R The scaler 30 obtains the input current frame I as needed. C and the number of pixels of the motion vector MV is expanded, and the input current frame I C ' and motion vector MV' are obtained (step S1).

[0070] Specialization coefficient database DBC and conversion parameter database DB P is the rendering information S R It is determined whether the rendering conditions have changed based on the above (step S2).

[0071] If the rendering conditions have not changed (step S2: No), C is the specialization coefficient D of DNN40 R The motion compensation unit 60 does not switch between the inference history frame I, which is the DNN output of the previous frame, based on the motion vector MV'. H (intermediate feature) and perform motion compensation to obtain the corrected history frame I Hc MC (Step S3).

[0072] The motion compensation unit 60 calculates the input current frame I C ' and the correction history frame I Hc MC is input to the trained DNN 40 (step S4). The DNN 40 outputs the inference result, O The DNN 40 outputs the intermediate feature obtained by the inference to the output unit 70 as the final output. H (Step S5) After that, the recursive processing of the RNN is repeated.

[0073] If the rendering conditions have changed (step S2: Yes), the conversion parameter database DB P is a conversion parameter P associated with the rendering conditions before and after the change. T The feature transformation unit 50 extracts the extracted transformation parameters P T The feature transform unit 50 obtains the transformation parameters P T , and the inference history frame I, which is the DNN output from the previous frame, is used. H (intermediate feature) to the input current frame I C The corrected feature quantity I corresponding to the rendering condition Hc (step S7).

[0074] Specialization coefficient database DB C is the input current frame I C Specialization coefficient D corresponding to the rendering conditions R is extracted and applied to the DNN 40 (step S8). Then, the process proceeds to step S3, and the corrected feature quantity I Hc (step S3) and inference processing by the DNN 40 (steps S4 and S5) are performed. Then, the recursive processing of the RNN is repeated until inference processing for all video frames is completed.

[0075] FIG. 13 shows the transformation parameters P T FIG. 10 is a diagram illustrating an example of a processing flow for calculating

[0076] The renderer 20 sets a plurality of rendering conditions to be specialized (step S11). The scaler 30 receives the input current frame I from the renderer 20. C , the motion vector MV and the rendering information S R The scaler 30 obtains the input current frame I as needed. C and the number of pixels of the motion vector MV is expanded, and the input current frame I C and motion vector MV' are obtained (step S12).

[0077] The motion compensation unit 60 calculates the inference history frame I, which is the DNN output from the previous frame, based on the motion vector MV'. H (intermediate feature) and perform motion compensation to obtain the corrected history frame I Hc MC (Step S13). The motion compensation unit 60 obtains the input current frame I C ' and the correction history frame I Hc MC The feature storage unit 90 inputs the intermediate feature I output from the DNN 40 (step S14). H DR is acquired as sample data and stored in association with the rendering conditions (step S15).

[0078] The feature amount storage unit 90 determines whether sample data has been acquired for all of the expected rendering conditions (step S16). If there is a rendering condition for which sample data has not been acquired (step S16: No), the process returns to step S11, and the above-described process is repeated until sample data has been acquired for all of the rendering conditions.

[0079] If sample data has been acquired for all the assumed rendering conditions (step S16: Yes), the conversion parameter P for converting between the sample data is calculated using the method described with reference to FIG. 9 or FIG. 10. T (Step S17). P is the transformation parameter P T is stored in association with the rendering conditions at the time of conversion (step S18).

[0080] 9. Example of Hardware Configuration FIG. 14 is a diagram illustrating an example of the hardware configuration of the information processing device 1. As shown in FIG.

[0081] The information processing of the information processing device 1 is realized by, for example, a computer 1000. The computer 1000 has a CPU (Central Processing Unit) 1100, a RAM (Random Access Memory) 1200, a ROM (Read Only Memory) 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.

[0082] The CPU 1100 operates and controls each component based on a program (program data 1450) stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the program stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0083] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 starts up, as well as programs that depend on the hardware of the computer 1000 .

[0084] The HDD 1400 is a non-transitory computer-readable recording medium that non-temporarily records programs executed by the CPU 1100 and data used by such programs. Specifically, the HDD 1400 is a recording medium that records an information processing program according to an embodiment as an example of program data 1450.

[0085] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0086] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display device, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.

[0087] For example, when the computer 1000 functions as the information processing device 1 according to the embodiment, the CPU 1100 of the computer 1000 executes an information processing program loaded onto the RAM 1200 to realize the functions of the aforementioned components. The information processing program, various models, and various data according to the present disclosure are stored in the HDD 1400. The CPU 1100 reads and executes program data 1450 from the HDD 1400. Alternatively, the CPU 1100 may acquire these programs from another device via an external network 1550.

[0088] [10. Effect] The information processing device 1 stores the specialization coefficient database DB C , a feature conversion unit 50, a motion compensation unit 60, and a DNN 40. C is the input current frame I C The specialization coefficient D is adjusted to match the change in rendering conditions to R The feature conversion unit 50 changes the specialization coefficient D before the rendering condition is changed. R The intermediate feature quantity for the past frame inferred using R The corrected feature I is estimated to be inferred by Hc The motion compensation unit 60 converts the corrected feature quantity I Hc , and perform motion compensation to obtain the corrected history frame I Hc MC The DNN 40 obtains the corrected history frame I Hc MC and the input current frame I C Based on the input current frame I C In the information processing method of the present disclosure, the processing of the information processing device 1 is executed by a computer 1000. The computer-readable non-transitory storage medium of the present disclosure stores a program that causes the computer 1000 to realize the processing of the information processing device 1.

[0089] According to this configuration, the specialization coefficient D R The intermediate feature amount is appropriately corrected in accordance with the change of the specialization coefficient D RAfter switching, the DNN 40 has a specialization coefficient D R Corrected feature I corresponding to the intermediate feature during learning Hc Therefore, the inference accuracy of the DNN 40 is less likely to decrease when the rendering conditions are switched.

[0090] The information processing device 1 includes a renderer 20, a feature amount storage unit 90, a conversion parameter calculation unit 80, and a conversion parameter database DB P The renderer 20 renders the same video scene for learning for each rendering condition. The feature storage unit 90 stores intermediate feature values ​​I H DR The transformation parameter calculation unit 80 calculates the intermediate feature I H DR A conversion model CV that converts between T Conversion parameter database DB P is the transformation parameter P T and the transformation parameter P T The two rendering conditions to be converted by are stored in association with each other.

[0091] According to this configuration, the intermediate feature is converted into the corrected feature I Hc A highly accurate conversion parameter P T is obtained.

[0092] Specialization coefficient database DB C acquires information about the rendering resolution as a rendering condition. C is the specialization coefficient D corresponding to the switching of the rendering resolution from the first resolution to the second resolution. R is switched from the first specialization coefficient to the second specialization coefficient.

[0093] According to this configuration, the specialization coefficient D R You can switch between the following.

[0094] The feature conversion unit 50 converts the intermediate feature obtained by applying the first resolution image to the first DNN into a corrected feature I that is estimated to be obtained when the second resolution image is applied to the second DNN. Hc Here, the video scene of the past frame is referred to as the past scene. The image obtained by rendering the past scene at the first resolution is referred to as the first resolution image. The image obtained by rendering the past scene at the second resolution is referred to as the second resolution image. The DNN to which the first specialization coefficient is applied is referred to as the first DNN. The DNN to which the second specialization coefficient is applied is referred to as the second DNN.

[0095] According to this configuration, appropriate conversion of intermediate feature amounts is performed in response to switching of rendering resolution.

[0096] The information processing device 1 includes a scaler 30. The scaler 30 scales the input current frame I C Scale the

[0097] According to this configuration, the DNN 40 can obtain highly accurate inference results.

[0098] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0099] [Additional Notes] The present technology may also be configured as follows: (1) An information processing device including: a specialized coefficient database that switches specialized coefficients in accordance with switching of rendering conditions from a past frame to an input current frame; a feature conversion unit that converts intermediate features related to the past frame that were inferred using the specialized coefficients before the switching of the rendering conditions into corrected features that are estimated to be inferred using the specialized coefficients after the switching of the rendering conditions; a motion compensation unit that performs motion compensation on the corrected features to obtain corrected history frames; and a DNN that performs inference related to the input current frame based on the corrected history frames and the input current frame. (2) The information processing device according to (1), comprising: a renderer that renders the same video scene for learning for each of the rendering conditions, a feature storage unit that stores the intermediate feature corresponding to each of the rendering conditions, a conversion parameter calculation unit that calculates a conversion model that converts between the intermediate feature of different rendering conditions and acquires conversion parameters to be applied to the conversion model, and a conversion parameter database that stores the conversion parameters in association with two rendering conditions that are to be converted by the conversion parameters. (3) The information processing device according to (1) or (2), wherein the specialization coefficient database acquires information on rendering resolution as the rendering condition, and switches the specialization coefficient from a first specialization coefficient to a second specialization coefficient in response to switching of the rendering resolution from a first resolution to a second resolution. (4) The information processing device described in (3) above, wherein when the video scene of the past frame is defined as a past scene, the image obtained by rendering the past scene at the first resolution is defined as a first resolution image, the image obtained by rendering the past scene at the second resolution is defined as a second resolution image, the DNN to which the first specialization coefficient is applied is defined as a first DNN, and the DNN to which the second specialization coefficient is applied is defined as a second DNN, the feature conversion unit converts the intermediate feature obtained by applying the first resolution image to the first DNN into the corrected feature estimated to be obtained when the second resolution image is applied to the second DNN.(5) The information processing device according to (4), further comprising a scaler that scales the input current frame in accordance with an architecture of the DNN. (6) An information processing method executed by a computer, comprising: switching specialized coefficients in accordance with switching of rendering conditions from a past frame to an input current frame; converting intermediate features related to the past frame that were inferred using the specialized coefficients before the switching of the rendering conditions into corrected features that are estimated to be inferred using the specialized coefficients after the switching of the rendering conditions; performing motion compensation on the corrected features to obtain corrected history frames; and performing inference related to the input current frame based on the corrected history frames and the input current frame. (7) A computer-readable non-transitory storage medium storing a program that causes a computer to perform the following: switching specialization coefficients in accordance with a switch in rendering conditions from a past frame to an input current frame; converting intermediate features related to the past frame that were inferred using the specialization coefficients before the switching of the rendering conditions into corrected features that are estimated to be inferred using the specialization coefficients after the switching of the rendering conditions; performing motion compensation on the corrected features to obtain corrected history frames; and performing inference related to the input current frame based on the corrected history frames and the input current frame.

[0100] REFERENCE SIGNS LIST 1 Information processing device 20 Renderer 30 Scaler 40 DNN 50 Feature conversion unit 60 Motion compensation unit 80 Conversion parameter calculation unit 90 Feature storage unit CV Conversion model DB C Specialization coefficient database DB P Conversion parameter database D R Specialization Coefficient I C Input current frame I Hc Corrected feature I Hc MC Correction History Frame P T Conversion parameters

Claims

1. An information processing device having: a specialized coefficient database that switches specialized coefficients in accordance with a change in rendering conditions from a past frame to an input current frame; a feature conversion unit that converts intermediate features related to the past frame inferred using the specialized coefficients before the rendering conditions are changed into corrected features estimated to be inferred using the specialized coefficients after the rendering conditions are changed; a motion compensation unit that performs motion compensation on the corrected features to obtain a corrected history frame; and a DNN that makes inferences related to the input current frame based on the corrected history frame and the input current frame.

2. An information processing device as described in claim 1, comprising: a renderer that renders the same video scene for learning for each of the rendering conditions; a feature storage unit that stores the intermediate features corresponding to each of the rendering conditions; a conversion parameter calculation unit that calculates a conversion model that converts between the intermediate features of the different rendering conditions and obtains conversion parameters to be applied to the conversion model; and a conversion parameter database that stores, in association with each other, the conversion parameters and two rendering conditions that are to be converted using the conversion parameters.

3. The information processing device according to claim 1, wherein the specialization coefficient database acquires information regarding rendering resolution as the rendering condition, and switches the specialization coefficient from a first specialization coefficient to a second specialization coefficient in response to the rendering resolution being switched from a first resolution to a second resolution.

4. The information processing device of claim 3, wherein when the video scene of the past frame is a past scene, an image obtained by rendering the past scene at the first resolution is a first resolution image, an image obtained by rendering the past scene at the second resolution is a second resolution image, the DNN to which the first specialization coefficient is applied is a first DNN, and the DNN to which the second specialization coefficient is applied is a second DNN, the feature conversion unit converts the intermediate feature obtained by applying the first resolution image to the first DNN into the corrected feature estimated to be obtained when the second resolution image is applied to the second DNN.

5. The information processing device according to claim 4, further comprising a scaler that scales the input current frame in accordance with an architecture of the DNN.

6. An information processing method executed by a computer, comprising: switching a specialization coefficient in accordance with a change in rendering conditions from a past frame to an input current frame; converting intermediate features related to the past frame inferred using the specialization coefficient before the rendering conditions were changed into corrected features estimated to be inferred using the specialization coefficient after the rendering conditions were changed; performing motion compensation on the corrected features to obtain a corrected history frame; and making an inference related to the input current frame based on the corrected history frame and the input current frame.

7. A computer-readable non-transitory storage medium storing a program that causes a computer to perform the following operations: switching specialization coefficients in accordance with a switch in rendering conditions from a past frame to an input current frame; converting intermediate features related to the past frame inferred using the specialization coefficients before the rendering conditions were switched into corrected features estimated to be inferred using the specialization coefficients after the rendering conditions were switched; performing motion compensation on the corrected features to obtain a corrected history frame; and performing inference related to the input current frame based on the corrected history frame and the input current frame.

Citation Information

Patent Citations

  • Game program and recording medium

    JP2018061674A

  • Image upsampling using one or more neural networks

    JP2023544231A